Back to Directory
NVIDIA Triton Inference Server logo

NVIDIA Triton Inference Server

2 use cases using this technology

Aerospace & DefenseGenerative AILarge Language ModelsRetrieval-Augmented GenerationAI Model Development & MLOps

Boosting Innovation and Cutting Costs Through Lockheed Martin's AI Factory

Lockheed Martin

Lockheed Martin centralized compute resources, MLOps tools and best practices into the Lockheed Martin AI Factory, built on an NVIDIA DGX SuperPOD reference architecture, to build and deploy trustworthy AI at scale on-premises under strict data governance requirements. Developers can now get GPU-backed environments running in minutes instead of weeks, and training times dropped from weeks to days. The AI factory processes over one billion tokens per week and now serves 7,000 engineers and developers, supporting internal chatbots and coding assistants such as Lockheed Martin Text Navigator, and has consolidated 30+ models on-premises.

LogisticsLarge Language Models

Delhivery achieves 160 ms latency for high-precision geocoding using Amazon EKS

Delhivery

Delhivery, a logistics provider in India, implemented a fine-tuned open-source Llama 3.2 1B large language model on Amazon EKS to support high-volume geocoding of pickup and drop-off addresses. The system processes up to 8,000 requests per minute at 160 milliseconds latency using NVIDIA A10G GPU-backed G5 Xlarge instances and the vLLM framework. Delhivery cut model-serving costs by approximately 80 percent and accelerated prototyping cycles from two days to under six hours, working with the AWS Prototyping and Cloud Engineering (PACE) team.