Grammarly migrates to a Databricks lakehouse to scale analytics for 30 million users
Grammarly replaced its homegrown analytics platform and always-on Amazon EMR clusters with the Databricks Data + AI Platform, using Delta Lake as the core of its lakehouse and Unity Catalog for fine-grained access control and data lineage, to consolidate 5 billion daily user-feedback events for analytics, achieving 110% faster querying at 10% of the ingestion cost of a data warehouse and cutting the time to make daily events available for analytics from 4 hours to under 15 minutes.
Overview
Grammarly replaced its homegrown analytics platform and always-on Amazon EMR clusters with the Databricks Data + AI Platform, using Delta Lake as the core of its lakehouse and Unity Catalog for fine-grained access control and data lineage, to consolidate 5 billion daily user-feedback events for analytics, achieving 110% faster querying at 10% of the ingestion cost of a data warehouse and cutting the time to make daily events available for analytics from 4 hours to under 15 minutes.
This entry has 13 published fields tied to exact passages in an immutable source capture.
Inspect the highlighted sourceThe challenge
Grammarly's homegrown legacy analytics platform, built on an in-house SQL-like language, was time-intensive to learn, could not effectively ingest external data or support Tableau reporting, ran costly always-on Amazon EMR clusters, and produced data silos as teams built their own analytics tools independently.
The solution
Grammarly migrated to the Databricks Data + AI Platform, using Delta Lake as the core of a central lakehouse and Unity Catalog for fine-grained, role-based access control and end-to-end data lineage, consolidating internal product event streams with external advertising data into a single source of truth queried via Databricks SQL and visualized in Tableau.
Reported business value
Grammarly achieved 110% faster querying at 10% of the ingestion cost of a data warehouse, and cut the time to make its 5 billion daily events available for analytics from 4 hours to under 15 minutes, helping marketing, sales and customer success teams get faster, more confident insights.
Sources
Open any source and check the claim yourself — that is the point of the register.
This record was researched and written with AI assistance, and its claims were checked against the sources above. (EU AI Act art. 50 transparency notice.)
Other technology & software entries in the register.
HP crafts marketing campaigns that resonate with customers using Databricks and Uniphore
HP centralized first-party customer data on the Databricks Data + AI Platform with Delta Lake and Unity Catalog, and connected it to Uniphore's HybridCompute for federated query pushdown, cutting campaign setup from 2 weeks to 2 hours and processing 400 million records in seconds.
Building a safer and more sustainable world
Novade, a construction management software company, partnered with Databricks to modernize its analytics and ML infrastructure, moving from a homegrown Apache Airflow stack to the Databricks Data + AI Platform with Delta Lake, Unity Catalog and MLflow. The move reduced total cost of ownership by 60% and supported a 100% increase in new clients, while enabling ML models for predicting incident risk affecting worker safety and project delivery schedules.
SafeGraph optimizes geospatial data processing with Databricks and Delta Sharing
SafeGraph used the Databricks Data + AI Platform, Delta Lake, Delta Sharing and MLflow to process petabytes of geospatial data and feed predictive models, achieving 2x-10x faster spatial querying, indexing and partitioning and a 50% reduction in peak memory consumption compared to other platforms, while reducing data access time for partners from months to minutes via Delta Sharing.
Transforming Weather Forecasting with Lakeflow Jobs
AccuWeather migrated from on-premises infrastructure to Databricks and Lakeflow Jobs, working with Datadog for observability, to unify diverse weather data formats and orchestrate 4,500+ weekly jobs. Lakeflow Jobs coordinates the ingestion of multiple weather models, triggers machine learning processes that weight and blend different forecasts, and manages complex job dependencies for reinforcement training workflows used in AccuWeather's proprietary forecasting engine. AccuWeather reports 3x faster dataset development (three months to one month per dataset), a 50% reduction in unactionable alerts, and 50% cost savings on serverless job usage.
Was this helpful?
Your feedback helps us improve our use case database
