From Data Chaos to a Trust Score for Every Table
Databricks' own internal Data Platform team rebuilt data governance across its internal lakehouse of more than 100,000 tables spanning engineering, go-to-market, HR, finance and product telemetry, using Unity Catalog to move from inconsistent, team-by-team access patterns to automatic, measurable governance. The team built a universal classification model tagging every table by sensitivity and domain, an internal system called Fortress enforcing purpose-bound, time-limited access via Unity Catalog's security APIs, and a continuous Data Governance Score (0-100) per dataset covering documentation, reliability and governance, propagated automatically through data lineage. The resulting governance layer now underpins AI agents used across every function at Databricks, from engineering to finance to facilities, letting teams check a dataset's governance score before using it to train a model instead of running manual reviews.
Overview
Databricks' own internal Data Platform team rebuilt data governance across its internal lakehouse of more than 100,000 tables spanning engineering, go-to-market, HR, finance and product telemetry, using Unity Catalog to move from inconsistent, team-by-team access patterns to automatic, measurable governance. The team built a universal classification model tagging every table by sensitivity and domain, an internal system called Fortress enforcing purpose-bound, time-limited access via Unity Catalog's security APIs, and a continuous Data Governance Score (0-100) per dataset covering documentation, reliability and governance, propagated automatically through data lineage. The resulting governance layer now underpins AI agents used across every function at Databricks, from engineering to finance to facilities, letting teams check a dataset's governance score before using it to train a model instead of running manual reviews.
This entry has 10 published fields tied to exact passages in an immutable source capture.
Inspect the highlighted sourceThe challenge
Before Unity Catalog, governance was inconsistent: access patterns varied by team, classifications were incomplete, and no one could confidently say which datasets were safe for AI training. Security, legal and engineering teams spent hours chasing spreadsheets to answer basic questions during compliance reviews, and new AI projects stalled because no one could confirm whether a dataset met privacy requirements.
The solution
Reported business value
Over 100,000 tables are now classified and governed automatically, eliminating manual audit cycles, with the Data Governance Score giving every dataset a measurable trust rating tied to AI readiness, used daily by more than 10,000 Databricks employees across every function.
Sources
Open any source and check the claim yourself — that is the point of the register.
This record was researched and written with AI assistance, and its claims were checked against the sources above. (EU AI Act art. 50 transparency notice.)
Other technology & software entries in the register.
Transforming Weather Forecasting with Lakeflow Jobs
AccuWeather migrated from on-premises infrastructure to Databricks and Lakeflow Jobs, working with Datadog for observability, to unify diverse weather data formats and orchestrate 4,500+ weekly jobs. Lakeflow Jobs coordinates the ingestion of multiple weather models, triggers machine learning processes that weight and blend different forecasts, and manages complex job dependencies for reinforcement training workflows used in AccuWeather's proprietary forecasting engine. AccuWeather reports 3x faster dataset development (three months to one month per dataset), a 50% reduction in unactionable alerts, and 50% cost savings on serverless job usage.
Supermetrics: Helping Marketers Redefine Efficiency with AI-Powered Data Analysis
Supermetrics, a Finland-based marketing intelligence platform serving 15,000+ customers across 132 countries, built an AI agent on Google Cloud using Vertex AI Agent Builder and the Agent Development Kit (ADK) that autonomously manages data connections, fixes pipeline errors, and analyzes campaign performance in real time, suggesting new creative options using Imagen. The agent automates the weekly marketing reporting cycle that previously took performance marketers up to four hours, reclaiming over 15 hours per month per marketer for strategy and creative testing. The system uses a central AI agent that interprets natural language requests and delegates tasks to sub-agents, and stores 'core memories' of user preferences for personalized context.
Domyn builds Colosseum 355B, a sovereign AI foundation model, using NVIDIA DGX Cloud
Domyn (formerly iGenius), an Italian AI company serving highly regulated sectors such as financial services and public administration, used NVIDIA DGX Cloud with over 3,000 NVIDIA H100 GPUs to continue-pretrain Colosseum 355B, a 355-billion-parameter foundation LLM. Within one week Domyn had access to the dedicated infrastructure, and within two months completed continued pretraining, achieving 82.04% accuracy on the MMLU benchmark. The model powers Domyn's business intelligence agent, Crystal, a sovereign AI solution deployed on private infrastructure.
Strava's Athlete Intelligence Translates Workout Data into Simple and Personalized Insights
Strava launched Athlete Intelligence, an AI-powered feature available as a public beta to subscribers, which analyzes and interprets workout data across pace, heart rate, elevation, power, and Relative Effort into simple, personalized insights and guidance. The feature spots 30-day performance trends, detects milestones such as fastest pace or longest distance, and offers tailored feedback for each activity, drawing on more than 10 billion activity uploads on Strava.
Was this helpful?
Your feedback helps us improve our use case database


