ElevenLabs: Scaling Voice AI with GPU-Accelerated Infrastructure
ElevenLabs
ElevenLabs runs its voice AI platform for text-to-speech, voice cloning and speech-to-speech translation on Google Kubernetes Engine (GKE) with NVIDIA GPUs, using GPU optimization strategies such as Multi-Instance GPUs and GPU time sharing to improve utilization and reduce costs. The platform generates 600 hours of audio for every hour of real time, supports 29 languages (expanding to 40), and is reportedly used by 41% of Fortune 500 companies, powering products including AI Dubbing Studio, long-form audio for audiobooks and podcasts, article narration for outlets like The New York Times, and real-time conversational AI assistants.
Overview
ElevenLabs runs its voice AI platform for text-to-speech, voice cloning and speech-to-speech translation on Google Kubernetes Engine (GKE) with NVIDIA GPUs, using GPU optimization strategies such as Multi-Instance GPUs and GPU time sharing to improve utilization and reduce costs. The platform generates 600 hours of audio for every hour of real time, supports 29 languages (expanding to 40), and is reportedly used by 41% of Fortune 500 companies, powering products including AI Dubbing Studio, long-form audio for audiobooks and podcasts, article narration for outlets like The New York Times, and real-time conversational AI assistants.
This entry has 12 published fields tied to exact passages in an immutable source capture.
Inspect the highlighted sourceThe challenge
As one of the fastest-growing AI companies, ElevenLabs needed to run massive-scale voice AI inference — producing 600 hours of generated audio for every hour of real time — while improving GPU utilization and controlling infrastructure costs.
The solution
ElevenLabs built its voice AI platform for text-to-speech, voice cloning and speech-to-speech translation on Google Kubernetes Engine (GKE) with NVIDIA GPUs, including H100s. It used GPU optimization strategies such as Multi-Instance GPUs (dividing a single GPU into seven hardware-isolated slices) and GPU Time Sharing (context-switching between processes on a GPU) to improve utilization and reduce costs. Its foundational model development followed an iterative approach: starting with a text-to-speech model, extending to multiple languages, adding voice cloning, and then building a speech-to-speech model for real-time translation.
Reported business value
The platform generates 600 hours of audio for every hour of real time, supports 29 languages (expanding to 40), and is reportedly used by 41% of Fortune 500 companies. It powers products including AI Dubbing Studio, long-form audio for audiobooks and podcasts, article narration for outlets like The New York Times, and real-time conversational AI assistants.
Sources
Open any source and check the claim yourself — that is the point of the register.
Other media & entertainment entries in the register.
How Business Insider's AI-based paywall strategy increased conversions by 75%
Business Insider replaced its editorially-driven freemium paywall with a machine-learning-based smart paywall that decides which content to paywall for each user based on prior reading habits, referral platform, and each content genre's propensity to convert. In testing from December to April, total conversions increased more than 75% versus the control, with 60% of new conversions coming from stories that would never have been paywalled before. Applying AI to the registration wall as well increased registrations by 300%.
Süddeutsche Zeitung: Lokalinformationen per WhatsApp
Süddeutsche Zeitung (SZ) launched WhatsApp channels for districts around Munich, offering free local news, event information and regional updates. Content comes from SZ's local newsrooms and is supplemented by automatically generated information produced in cooperation with Berlin media startup Beat Squares, which uses AI to process data from publicly accessible sources. All contributions are editorially reviewed before publication. More than 4,000 users subscribed within the first three days.
Wiley boosts service efficiency with Salesforce Agentforce
Publisher Wiley uses Salesforce Agentforce alongside Service Cloud and Einstein AI to manage customer service, letting AI agents resolve common issues like password resets and account access so human reps can focus on complex cases. Agentforce improved case resolution by over 40% compared to Wiley's previous chatbot, helped onboard seasonal agents 50% faster, and delivered a 213% return on investment with $230,000 in annual cost savings.
From lab to newsroom: How Reuters builds AI tools journalists actually use
Reuters developed a suite of AI-powered newsroom tools with a human-in-the-loop approach: Fact Genie, an AI-assisted summarisation tool that scans entire documents in under five seconds to suggest newsworthy alerts (helping Speed teams publish the first alert within about six seconds of a press release, across roughly 100,000 business news alerts published monthly by 250-300 journalists); LEON, an AI-powered headline assistant; and AVISTA, a machine-learning tool for sourcing, tagging and archiving photos and videos. Fact Genie took about four months from prototype to production rollout, starting with senior journalists in the US and UK before expanding, and the team also uses AI to filter out non-newsworthy content before it reaches a large language model, cutting alert-generation time from about a minute with GPT-3.5 Turbo to about 10 seconds with GPT-4o mini.
Was this helpful?
Your feedback helps us improve our use case database