vLLM
2 use cases using this technology
RTLZWEI Hones Competitive Edge with Red Hat OpenShift AI
RTLZWEI
RTLZWEI, a German broadcaster and digital media company, modernized its on-premise core system BOBY by deploying Red Hat OpenShift Platform Plus and Red Hat OpenShift AI, with guidance from Red Hat Consulting and Red Hat Training. A proof of concept integrated NVIDIA H100 GPUs with Red Hat OpenShift AI, using Whisper on vLLM to speed up video transcription for subtitles, dubbing and accessibility. By fine-tuning the transcription model with its own data, RTLZWEI reduced the word error rate by 33%. The platform also hosts LLMs for data science experiments, forecasting and fine-tuning, running on-premise to meet data sovereignty and cybersecurity requirements.
Delhivery achieves 160 ms latency for high-precision geocoding using Amazon EKS
Delhivery
Delhivery, a logistics provider in India, implemented a fine-tuned open-source Llama 3.2 1B large language model on Amazon EKS to support high-volume geocoding of pickup and drop-off addresses. The system processes up to 8,000 requests per minute at 160 milliseconds latency using NVIDIA A10G GPU-backed G5 Xlarge instances and the vLLM framework. Delhivery cut model-serving costs by approximately 80 percent and accelerated prototyping cycles from two days to under six hours, working with the AWS Prototyping and Cloud Engineering (PACE) team.