Kythera Labs builds its business on Spark Declarative Pipelines and the Databricks Data + AI Platform
Kythera Labs, a healthcare big-data company processing about 2 billion medical and 3 billion prescription transactions annually, uses the Databricks Data + AI Platform with Spark Declarative Pipelines, Delta Sharing and Unity Catalog to remaster over 3 petabytes of raw healthcare data into analysis-ready products via its Wayfinder offering. This delivered a 99%+ reduction in query pipeline processing time (from 2 days to 2-4 minutes) and 80% faster time to market.
Overview
Kythera Labs, a healthcare big-data company processing about 2 billion medical and 3 billion prescription transactions annually, uses the Databricks Data + AI Platform with Spark Declarative Pipelines, Delta Sharing and Unity Catalog to remaster over 3 petabytes of raw healthcare data into analysis-ready products via its Wayfinder offering. This delivered a 99%+ reduction in query pipeline processing time (from 2 days to 2-4 minutes) and 80% faster time to market.
This entry has 11 published fields tied to exact passages in an immutable source capture.
Inspect the highlighted sourceThe challenge
Healthcare and life sciences organizations often struggle to extract actionable insights and value from the massive and continuously increasing amount of complex data available. When Kythera was founded in 2019, the company struggled to store and process its big healthcare data efficiently and cost-effectively.
The solution
Kythera uses Spark Declarative Pipelines to automate and simplify the orchestration of ETL pipelines, which helps deliver better insights because the data is more reliable and higher quality. The lakehouse architecture unifies data so Kythera can use it for both analysis and machine learning while reducing the risk of data egress and helping customers lower costs.
Sources
Open any source and check the claim yourself — that is the point of the register.
This record was researched and written with AI assistance, and its claims were checked against the sources above. (EU AI Act art. 50 transparency notice.)
Other healthcare entries in the register.
Less paperwork, more time for people: Azure OpenAI at Landsberg Hospital
Klinikum Landsberg am Lech, a German acute-care hospital, deployed an AI language solution built on Azure OpenAI with partner Pexon Consulting that transcribes doctor-patient consultations in real time via tablet, recognizes medical terms, structures the content, and transfers it directly into the hospital information system.
Novo Nordisk builds AI drug discovery platform on Azure with Microsoft Research
Novo Nordisk partnered with Microsoft Research to build an AI platform on Azure AI and data stacks spanning regulatory affairs, early research, drug discovery and trial design, using Azure OpenAI Service, Azure Cosmos DB and Azure Kubernetes Service, with Power BI and Power Apps for collaboration. The platform includes a copilot for researchers, shared reasoning-chain templates, and governance/auditing of how data and models are used. The teams published early results on predictive AI models for cardiovascular disease risk detection, including an algorithm that Novo Nordisk says predicts patients' cardiovascular risk better than the best clinical standards, drawing on more than 100 years of insulin research data.
BAYADA Builds a Unified, AI-Ready Platform for Compassionate Care
BAYADA Home Health Care is consolidating three separate legacy data platforms (multiple practice management systems, an on-prem ODS, and Snowflake) and 65+ enterprise data sources into a single Databricks Lakehouse under its Data Modernization program, using medallion tiers, Lakeflow Jobs, Asset Bundles, and Unity Catalog governance. As part of the migration, BAYADA uses an LLM-powered code converter and Databricks Assistant to automate SQL and stored-procedure translation, and applies a machine learning-based data mastering accelerator to create golden Client, Payor, Candidate, and Referrer records. BAYADA reports roughly 30% better workload performance and cost efficiency versus its prior Snowflake environment, and is laying a governed foundation for AI agents supporting payroll validation, compliance monitoring, and operational insights, separately from its roadmap for AI-assisted chart review and risk prediction.
Heidi Health fine-tunes NVIDIA Parakeet/Nemotron ASR, cutting clinical transcription latency 75% and costs 64%
Heidi Health, which processes 2.4 million clinical consultations weekly across 110 languages, fine-tuned the open NVIDIA Parakeet V2 speech recognition model on 1,500 hours of curated clinical audio using the NVIDIA NeMo framework and 8x H100 GPUs, replacing its closed-source vendor ASR. The in-house Nemotron-based ASR stack cut end-to-end transcription latency from about 3.0 seconds to 0.7 seconds (75%+ improvement), reduced word error rate from 13% to 9.4%, and cut operating costs by 64% by eliminating per-minute API pricing, while giving Heidi full ownership of model weights for data sovereignty.
Was this helpful?
Your feedback helps us improve our use case database

