Technology & SoftwareGenerative AIPublic CloudNVIDIA TensorRT-LLMNVIDIA A100 Tensor Core GPUsNVIDIA H100 Tensor Core GPUsAmazon EC2 P4dAmazon EC2 P5

Accelerating Large Language Model Inference with NVIDIA in the Cloud

Perplexity

Perplexity built pplx-api, an API for developers to integrate open-source LLMs with fast inference, served on Amazon EC2 P4d instances powered by NVIDIA A100 Tensor Core GPUs and accelerated with NVIDIA TensorRT-LLM (with a planned move to Amazon P5 instances with NVIDIA H100 GPUs). pplx-api achieves up to 3.1X lower latency and up to 4.3X lower first-token latency versus other deployment platforms, and switching external inference-serving API references to pplx-api lowered costs 4X, saving $600,000 per year. Using NVIDIA H100 GPUs and FP8 precision on Amazon P5 instances cuts latency in half and boosts throughput by 200 percent versus A100 GPUs in the same configuration. Perplexity also uses AWS's Kubernetes integration to scale elastically beyond hundreds of GPUs.

Overview

Perplexity built pplx-api, an API for developers to integrate open-source LLMs with fast inference, served on Amazon EC2 P4d instances powered by NVIDIA A100 Tensor Core GPUs and accelerated with NVIDIA TensorRT-LLM (with a planned move to Amazon P5 instances with NVIDIA H100 GPUs). pplx-api achieves up to 3.1X lower latency and up to 4.3X lower first-token latency versus other deployment platforms, and switching external inference-serving API references to pplx-api lowered costs 4X, saving $600,000 per year. Using NVIDIA H100 GPUs and FP8 precision on Amazon P5 instances cuts latency in half and boosts throughput by 200 percent versus A100 GPUs in the same configuration. Perplexity also uses AWS's Kubernetes integration to scale elastically beyond hundreds of GPUs.

The challenge

Delivering fast and efficient LLM inference is critical for real-time applications. As a startup, Perplexity faced escalating costs associated with LLM inference to support its rapid growth, while needing to maintain strict service-level agreement requirements and adapt quickly to an explosively growing ecosystem of community LLMs.

The solution

Perplexity built pplx-api, an API for developers to integrate open-source LLMs with fast inference, served on Amazon EC2 P4d instances powered by NVIDIA A100 Tensor Core GPUs and accelerated with NVIDIA TensorRT-LLM, with a planned full transition to Amazon P5 instances powered by NVIDIA H100 Tensor Core GPUs. Perplexity uses AWS's integration with Kubernetes to scale elastically beyond hundreds of GPUs.

Generative AILarge Language Models

Reported business value

pplx-api achieves up to 3.1X lower latency and up to 4.3X lower first-token latency relative to other deployment platforms. Switching external inference-serving API references to pplx-api lowered costs 4X, saving $600,000 per year. Using NVIDIA H100 GPUs and FP8 precision on Amazon P5 instances cuts latency in half and boosts throughput by 200 percent compared to NVIDIA A100 GPUs in the same configuration.

Sources

Open any source and check the claim yourself — that is the point of the register.

Related entries

Other technology & software entries in the register.

All entries
Technology & SoftwareAgentic AIPublic Cloud

Supermetrics: Helping Marketers Redefine Efficiency with AI-Powered Data Analysis

Supermetrics, a Finland-based marketing intelligence platform serving 15,000+ customers across 132 countries, built an AI agent on Google Cloud using Vertex AI Agent Builder and the Agent Development Kit (ADK) that autonomously manages data connections, fixes pipeline errors, and analyzes campaign performance in real time, suggesting new creative options using Imagen. The agent automates the weekly marketing reporting cycle that previously took performance marketers up to four hours, reclaiming over 15 hours per month per marketer for strategy and creative testing. The system uses a central AI agent that interprets natural language requests and delegates tasks to sub-agents, and stores 'core memories' of user preferences for personalized context.

96/100HighSupermetricsPrimary source
Technology & SoftwareLarge Language ModelsPublic Cloud

Domyn builds Colosseum 355B, a sovereign AI foundation model, using NVIDIA DGX Cloud

Domyn (formerly iGenius), an Italian AI company serving highly regulated sectors such as financial services and public administration, used NVIDIA DGX Cloud with over 3,000 NVIDIA H100 GPUs to continue-pretrain Colosseum 355B, a 355-billion-parameter foundation LLM. Within one week Domyn had access to the dedicated infrastructure, and within two months completed continued pretraining, achieving 82.04% accuracy on the MMLU benchmark. The model powers Domyn's business intelligence agent, Crystal, a sovereign AI solution deployed on private infrastructure.

96/100HighDomynPrimary source
Technology & SoftwareRecommendation & PersonalizationUnknown

Strava's Athlete Intelligence Translates Workout Data into Simple and Personalized Insights

Strava launched Athlete Intelligence, an AI-powered feature available as a public beta to subscribers, which analyzes and interprets workout data across pace, heart rate, elevation, power, and Relative Effort into simple, personalized insights and guidance. The feature spots 30-day performance trends, detects milestones such as fastest pace or longest distance, and offers tailored feedback for each activity, drawing on more than 10 billion activity uploads on Strava.

88/100HighStravaPrimary source
Technology & SoftwareMachine LearningPublic Cloud

Fifth Dimension unlocks insights and intelligence with Google Cloud

Fifth Dimension, founded in 2023, provides an AI platform combining machine learning, predictive analytics and natural language processing to uncover hidden patterns in real estate data, processing over 1TB of data per month by 2025. Facing scalability and cost problems with its original infrastructure, the company adopted Google Cloud, building its ML stack on Vertex AI (running Google Cloud Gemini and Anthropic Claude models), plus Cloud SQL, Cloud Run and Pub/Sub. This scaled processing capacity 50x to handle document-processing surges, supported 6x global client growth across multiple regions, and decreased infrastructure costs by 30% through serverless architecture. Model deployment time dropped from weeks to days, platform engagement tripled in 2025, and annual recurring revenue grew 6x in the same year.

96/100HighFifth DimensionPrimary source

Was this helpful?

Your feedback helps us improve our use case database