Achieving Near-Zero Downtime and Powering Generative AI Using Amazon EKS with Ada
Customer service automation company Ada migrated its Kubernetes clusters to Amazon EKS and now runs a generative AI reasoning engine for customer conversations, raising automated resolution rates from 20-30 percent with its prior declarative agents to up to 77 percent.
Overview
Customer service automation company Ada migrated its Kubernetes clusters to Amazon EKS and now runs a generative AI reasoning engine for customer conversations, raising automated resolution rates from 20-30 percent with its prior declarative agents to up to 77 percent.
This entry has 13 published fields tied to exact passages in an immutable source capture.
Inspect the highlighted sourceThe challenge
Ada's self-managed Kubernetes clusters (seven clusters with up to 700 nodes each) were complex and costly to maintain: two or three full-time engineers took up to 5 months to complete a Kubernetes version upgrade, and upgrades consumed up to 30% of the yearly error budget under a 99.9% availability SLA.
The solution
Ada migrated its self-managed Kubernetes clusters to Amazon EKS, using AWS Global Accelerator and a blue/green deployment strategy to incrementally route traffic with near-zero downtime, adopted GPU slicing for its ML inference workloads, used Argo CD to keep application sets in sync, and now runs a generative AI reasoning engine (using ML models and LLM calls) for its customer service automation product.
Reported business value
Ada cut compute costs 15%, increased compute efficiency 30%, increased GPU usage cost efficiency 20%, increased deployment velocity 70%, and cut Kubernetes upgrade time from up to 5 months to 5 days; its shift to a generative AI agent lifted automated resolution rates from 20-30% with declarative agents to up to 77%.
Sources
Open any source and check the claim yourself — that is the point of the register.
This record was researched and written with AI assistance, and its claims were checked against the sources above. (EU AI Act art. 50 transparency notice.)
Other technology & software entries in the register.
HP crafts marketing campaigns that resonate with customers using Databricks and Uniphore
HP centralized first-party customer data on the Databricks Data + AI Platform with Delta Lake and Unity Catalog, and connected it to Uniphore's HybridCompute for federated query pushdown, cutting campaign setup from 2 weeks to 2 hours and processing 400 million records in seconds.
Building a safer and more sustainable world
Novade, a construction management software company, partnered with Databricks to modernize its analytics and ML infrastructure, moving from a homegrown Apache Airflow stack to the Databricks Data + AI Platform with Delta Lake, Unity Catalog and MLflow. The move reduced total cost of ownership by 60% and supported a 100% increase in new clients, while enabling ML models for predicting incident risk affecting worker safety and project delivery schedules.
SafeGraph optimizes geospatial data processing with Databricks and Delta Sharing
SafeGraph used the Databricks Data + AI Platform, Delta Lake, Delta Sharing and MLflow to process petabytes of geospatial data and feed predictive models, achieving 2x-10x faster spatial querying, indexing and partitioning and a 50% reduction in peak memory consumption compared to other platforms, while reducing data access time for partners from months to minutes via Delta Sharing.
Transforming Weather Forecasting with Lakeflow Jobs
AccuWeather migrated from on-premises infrastructure to Databricks and Lakeflow Jobs, working with Datadog for observability, to unify diverse weather data formats and orchestrate 4,500+ weekly jobs. Lakeflow Jobs coordinates the ingestion of multiple weather models, triggers machine learning processes that weight and blend different forecasts, and manages complex job dependencies for reinforcement training workflows used in AccuWeather's proprietary forecasting engine. AccuWeather reports 3x faster dataset development (three months to one month per dataset), a 50% reduction in unactionable alerts, and 50% cost savings on serverless job usage.
Was this helpful?
Your feedback helps us improve our use case database
