{"slug":"accelerating-large-language-model-inference-with-nvidia-in-the-cloud","url":"https://findausecase.com/use-cases/accelerating-large-language-model-inference-with-nvidia-in-the-cloud","title":"Accelerating Large Language Model Inference with NVIDIA in the Cloud","description":"Perplexity built pplx-api, an API for developers to integrate open-source LLMs with fast inference, served on Amazon EC2 P4d instances powered by NVIDIA A100 Tensor Core GPUs and accelerated with NVIDIA TensorRT-LLM (with a planned move to Amazon P5 instances with NVIDIA H100 GPUs). pplx-api achieves up to 3.1X lower latency and up to 4.3X lower first-token latency versus other deployment platforms, and switching external inference-serving API references to pplx-api lowered costs 4X, saving $600,000 per year. Using NVIDIA H100 GPUs and FP8 precision on Amazon P5 instances cuts latency in half and boosts throughput by 200 percent versus A100 GPUs in the same configuration. Perplexity also uses AWS's Kubernetes integration to scale elastically beyond hundreds of GPUs.","company":"Perplexity","industry":"Technology & Software","aiCapabilities":["Generative AI","Large Language Models"],"technology":["NVIDIA TensorRT-LLM","NVIDIA A100 Tensor Core GPUs","NVIDIA H100 Tensor Core GPUs","Amazon EC2 P4d","Amazon EC2 P5"],"deployment":"Public Cloud","problemStatement":"Delivering fast and efficient LLM inference is critical for real-time applications. As a startup, Perplexity faced escalating costs associated with LLM inference to support its rapid growth, while needing to maintain strict service-level agreement requirements and adapt quickly to an explosively growing ecosystem of community LLMs.","solutionApproach":"Perplexity built pplx-api, an API for developers to integrate open-source LLMs with fast inference, served on Amazon EC2 P4d instances powered by NVIDIA A100 Tensor Core GPUs and accelerated with NVIDIA TensorRT-LLM, with a planned full transition to Amazon P5 instances powered by NVIDIA H100 Tensor Core GPUs. Perplexity uses AWS's integration with Kubernetes to scale elastically beyond hundreds of GPUs.","businessValue":"pplx-api achieves up to 3.1X lower latency and up to 4.3X lower first-token latency relative to other deployment platforms. Switching external inference-serving API references to pplx-api lowered costs 4X, saving $600,000 per year. Using NVIDIA H100 GPUs and FP8 precision on Amazon P5 instances cuts latency in half and boosts throughput by 200 percent compared to NVIDIA A100 GPUs in the same configuration.","evidence":{"band":"high"},"sourceUrl":"https://www.nvidia.com/en-us/case-studies/perplexity","dates":{"publishedAt":"2026-08-15T23:34:37.177Z","publishedAtSource":"ledger","updatedAt":"2026-08-18T10:29:43.045Z"},"license":"Open for reading and citing with a link to https://findausecase.com/use-cases/accelerating-large-language-model-inference-with-nvidia-in-the-cloud. Bulk republication requires permission."}