Media & EntertainmentSpeech & Audio AIGenerative AIConversational AI
ElevenLabs: Scaling Voice AI with GPU-Accelerated Infrastructure
ElevenLabs
ElevenLabs runs its voice AI platform for text-to-speech, voice cloning and speech-to-speech translation on Google Kubernetes Engine (GKE) with NVIDIA GPUs, using GPU optimization strategies such as Multi-Instance GPUs and GPU time sharing to improve utilization and reduce costs. The platform generates 600 hours of audio for every hour of real time, supports 29 languages (expanding to 40), and is reportedly used by 41% of Fortune 500 companies, powering products including AI Dubbing Studio, long-form audio for audiobooks and podcasts, article narration for outlets like The New York Times, and real-time conversational AI assistants.