
NVIDIA Triton Inference Server
2 use cases using this technology
Boosting Innovation and Cutting Costs Through Lockheed Martin's AI Factory
Lockheed Martin
Lockheed Martin centralized compute resources, MLOps tools and best practices into the Lockheed Martin AI Factory, built on an NVIDIA DGX SuperPOD reference architecture, to build and deploy trustworthy AI at scale on-premises under strict data governance requirements. Developers can now get GPU-backed environments running in minutes instead of weeks, and training times dropped from weeks to days. The AI factory processes over one billion tokens per week and now serves 7,000 engineers and developers, supporting internal chatbots and coding assistants such as Lockheed Martin Text Navigator, and has consolidated 30+ models on-premises.
Delhivery achieves 160 ms latency for high-precision geocoding using Amazon EKS
Delhivery
Delhivery, a logistics provider in India, implemented a fine-tuned open-source Llama 3.2 1B large language model on Amazon EKS to support high-volume geocoding of pickup and drop-off addresses. The system processes up to 8,000 requests per minute at 160 milliseconds latency using NVIDIA A10G GPU-backed G5 Xlarge instances and the vLLM framework. Delhivery cut model-serving costs by approximately 80 percent and accelerated prototyping cycles from two days to under six hours, working with the AWS Prototyping and Cloud Engineering (PACE) team.