Synthesia Accelerates AI Text-to-Video Model Training 30x on NVIDIA GPUs via AWS
AI video company Synthesia moved ML model training for its text-to-video SaaS, Synthesia Studio, from on-premises computers to NVIDIA GPU-powered Amazon EC2 instances (P5/H100, P4/A100, G5/A10G) using PyTorch, Amazon EKS, AWS ParallelCluster and AWS Batch, accelerating model training by 30 times and supporting 456% user growth.
Overview
AI video company Synthesia moved ML model training for its text-to-video SaaS, Synthesia Studio, from on-premises computers to NVIDIA GPU-powered Amazon EC2 instances (P5/H100, P4/A100, G5/A10G) using PyTorch, Amazon EKS, AWS ParallelCluster and AWS Batch, accelerating model training by 30 times and supporting 456% user growth.
This entry has 13 published fields tied to exact passages in an immutable source capture.
Inspect the highlighted sourceThe challenge
After growing rapidly to 350 employees and 50,000 customers, Synthesia found that training ML models on on-premises computers had become inefficient. Its production facilities generate many terabytes of data each week, and with more than 50 ML researchers training large models, the company needed a large, scalable data lake and compute cluster to run multiple generative AI models 24/7.
The solution
Synthesia switched to multi-node compute clusters for distributed ML model training on Amazon EC2 P5 Instances (NVIDIA H100 GPUs) and Amazon EC2 P4 Instances (NVIDIA A100 GPUs), and uses Amazon EC2 G5 Instances (NVIDIA A10G GPUs) for data processing and to optimize video rendering runtime. The company manages compute capacity using Amazon EKS, AWS ParallelCluster and AWS Batch, stores large datasets in Amazon S3, and builds its AI workflows on PyTorch and NVIDIA CUDA for video rendering and inference.
Reported business value
Using these managed services and NVIDIA GPU-powered compute instances, Synthesia reduced ML model training time for smaller voice models from days to hours, and supported a user base growth of 456 percent.
Sources
Open any source and check the claim yourself — that is the point of the register.
This record was researched and written with AI assistance, and its claims were checked against the sources above. (EU AI Act art. 50 transparency notice.)
Other technology & software entries in the register.
HP crafts marketing campaigns that resonate with customers using Databricks and Uniphore
HP centralized first-party customer data on the Databricks Data + AI Platform with Delta Lake and Unity Catalog, and connected it to Uniphore's HybridCompute for federated query pushdown, cutting campaign setup from 2 weeks to 2 hours and processing 400 million records in seconds.
Building a safer and more sustainable world
Novade, a construction management software company, partnered with Databricks to modernize its analytics and ML infrastructure, moving from a homegrown Apache Airflow stack to the Databricks Data + AI Platform with Delta Lake, Unity Catalog and MLflow. The move reduced total cost of ownership by 60% and supported a 100% increase in new clients, while enabling ML models for predicting incident risk affecting worker safety and project delivery schedules.
SafeGraph optimizes geospatial data processing with Databricks and Delta Sharing
SafeGraph used the Databricks Data + AI Platform, Delta Lake, Delta Sharing and MLflow to process petabytes of geospatial data and feed predictive models, achieving 2x-10x faster spatial querying, indexing and partitioning and a 50% reduction in peak memory consumption compared to other platforms, while reducing data access time for partners from months to minutes via Delta Sharing.
Transforming Weather Forecasting with Lakeflow Jobs
AccuWeather migrated from on-premises infrastructure to Databricks and Lakeflow Jobs, working with Datadog for observability, to unify diverse weather data formats and orchestrate 4,500+ weekly jobs. Lakeflow Jobs coordinates the ingestion of multiple weather models, triggers machine learning processes that weight and blend different forecasts, and manages complex job dependencies for reinforcement training workflows used in AccuWeather's proprietary forecasting engine. AccuWeather reports 3x faster dataset development (three months to one month per dataset), a 50% reduction in unactionable alerts, and 50% cost savings on serverless job usage.
Was this helpful?
Your feedback helps us improve our use case database
