← Back to Delhivery achieves 160 ms latency for high-precision geocoding using Amazon EKS

Source proof for Delhivery achieves 160 ms latency for high-precision geocoding using Amazon EKS

Source-bound proof

Verified source excerpts for every supported field

Each colour maps a published value to the exact source passage used to support it. Only bounded excerpts are public; administrators can inspect the complete captured source.

12 fields supported

Title

Delhivery achieves 160 ms latency for high-precision geocoding using Amazon EKS

derived · high
…ormation with expert guidance and packaged solutions AWS Solutions Case Studies Delhivery achieves 160 ms latency for high-precision geocoding using Amazon EKS Learn how Delhivery improved geocoding in logistics across India using generati…

Description

Delhivery, a logistics provider in India, implemented a fine-tuned open-source Llama 3.2 1B large language model on Amazon EKS to support high-volume geocoding of pickup and drop-off addresses. The system processes up to 8,000 requests per minute at 160 milliseconds latency using NVIDIA A10G GPU-backed G5 Xlarge instances and the vLLM framework. Delhivery cut model-serving costs by approximately 80 percent and accelerated prototyping cycles from two days to under six hours, working with the AWS Prototyping and Cloud Engineering (PACE) team.

derived · high
…azon Web Services (AWS) to support high-volume geocoding across its operations. The system now processes up to 8,000 requests per minute at 160 milliseconds latency, meeting internal benchmarks for real-time performance. As a result, Delhivery…

Company

Delhivery

quote · high
…w To improve address matching with greater speed, cost efficiency, and control, Delhivery , a logistics provider in India, turned to generative AI. The company implemented a fine-tuned large language m…

Country

India

classification · high
…ccelerated prototyping cycles from two days to under six hours. About Delhivery Delhivery is a technology-driven logistics and supply chain services provider in India, offering transportation, warehousing, freight, and fulfillment solutions to bu…

Industry

Logistics

classification · high
…ccelerated prototyping cycles from two days to under six hours. About Delhivery Delhivery is a technology-driven logistics and supply chain services provider in India, offering transportation, warehousing, freight, and fulfillment solutions to businesses nationwide. Opportunity | Using generative AI for high-precision…

Deployment model

Cloud

classification · high
…stance types, optimizing model serving with NVIDIA Triton Inference Server, and packaging the deployment for Amazon Elastic Kubernetes Service (Amazon EKS), which was already part of Delhivery’s tech stack. “We started with basic bench…

Problem

Delhivery initially tested serverless LLMs from third-party providers for high-precision geocoding of pickup and drop-off addresses, but these came with rate caps of 2,000 requests per minute or higher costs for provisioned access exceeding actual usage needs. Traditional machine learning models also lacked contextual understanding and required long training cycles, slowing experimentation and making it difficult to scale during demand spikes.

derived · high
…Delhivery team initially tested serverless LLMs from third-party providers. But these services came with limitations—rate caps of 2,000 requests per minute, or higher costs for provisioned access that exceeded actual usage needs. “We needed a solution that could handle up to 8,000 requests per minute and sti…

Solution

Delhivery selected and externally fine-tuned the open-source Llama 3.2 1B model, then engaged the AWS Prototyping and Cloud Engineering (PACE) team to identify suitable instance types, optimize model serving with NVIDIA Triton Inference Server, and package the deployment for Amazon EKS. Production deployment uses Amazon EKS with G5 Xlarge instances equipped with NVIDIA A10G GPUs and the vLLM framework, with auto scaling to handle demand.

derived · high
…de the transition to production much faster.” To support production deployment, the team used Amazon EKS with G5 Xlarge instances as cluster nodes. These instances, equipped with NVIDIA A10G GPUs, were chosen to deliver fast inference response times using the vLLM framework. With auto scaling in place, the Amazon EKS cluster can seamlessly scale the dep…

Technology

Llama 3.2 1B, Amazon Elastic Kubernetes Service, Amazon EC2 G5 Instances, NVIDIA A10G GPU, NVIDIA Triton Inference Server, vLLM

classification · high
…at could offer greater flexibility and performance. After testing various LLMs, Delhivery selected a version of the open-source Llama 3.2 1B model that delivered the performance it needed. The team fine-tuned the model externa…

Business value

The system processes up to 8,000 requests per minute at a latency of 160 milliseconds (measured at a concurrency of 30). Delhivery reduced its monthly model-serving costs by approximately 80 percent and accelerated prototyping cycles from two days to under six hours.

derived · high
…oughput required for Delhivery’s high-volume geocoding workloads. Additionally, Delhivery reduced its monthly model-serving costs by approximately 80 percent by optimizing GPU utilization and eliminating the overhead of third-party API provisioning. “By maximizing th…

Deployment options

cloud

classification · high
…stance types, optimizing model serving with NVIDIA Triton Inference Server, and packaging the deployment for Amazon Elastic Kubernetes Service (Amazon EKS), which was already part of Delhivery’s tech stack. “We started with basic bench…

AI capabilities

Large Language Models

classification · high
…at could offer greater flexibility and performance. After testing various LLMs, Delhivery selected a version of the open-source Llama 3.2 1B model that delivered the performance it needed. The team fine-tuned the model externally and began designing a solution tailore…
Capture details
Captured
14 Aug 2026, 09:35 UTC
Extractor
fetch-strip@1
Snapshot hash
5c96d0bff1384bdd69590fdff16529984657dcf9c842596f6db76da1c67db70c