Back to Directory
NVIDIA TensorRT-LLM logo

NVIDIA TensorRT-LLM

2 use cases using this technology

ManufacturingLarge Language ModelsGenerative AIAgentic AICode Generation

MediaTek Accelerates AI Development With an AI Factory

MediaTek

MediaTek established an on-premises AI factory powered by NVIDIA DGX SuperPOD with NVIDIA Blackwell-based systems to accelerate enterprise AI efforts, including development of its Breeze series LLMs and a 480-billion-parameter traditional-Chinese model. The AI factory processes approximately 60 billion tokens per month for inference and completes over 24,000 model-training iterations monthly, training models exceeding 480 billion parameters within one week (versus 7-billion-parameter models in a week previously). Using NVIDIA NIM and TensorRT-LLM, MediaTek achieved a 40% improvement in inference speed and 60% increase in token throughput. NVIDIA Mission Control consolidated GPU provisioning and system monitoring, while AI-assisted code completion and an AI agent for chip design documentation reduced documentation time from weeks to days. NVIDIA Riva was integrated into NVIDIA DGX Spark for agentic voice control features like internet search, calendar and messaging.

Technology & SoftwareGenerative AILarge Language Models

Accelerating Large Language Model Inference with NVIDIA in the Cloud

Perplexity

Perplexity built pplx-api, an API for developers to integrate open-source LLMs with fast inference, served on Amazon EC2 P4d instances powered by NVIDIA A100 Tensor Core GPUs and accelerated with NVIDIA TensorRT-LLM (with a planned move to Amazon P5 instances with NVIDIA H100 GPUs). pplx-api achieves up to 3.1X lower latency and up to 4.3X lower first-token latency versus other deployment platforms, and switching external inference-serving API references to pplx-api lowered costs 4X, saving $600,000 per year. Using NVIDIA H100 GPUs and FP8 precision on Amazon P5 instances cuts latency in half and boosts throughput by 200 percent versus A100 GPUs in the same configuration. Perplexity also uses AWS's Kubernetes integration to scale elastically beyond hundreds of GPUs.