{"slug":"delhivery-achieves-160-ms-latency-for-high-precision-geocoding-using-amazon-eks","url":"https://findausecase.com/use-cases/delhivery-achieves-160-ms-latency-for-high-precision-geocoding-using-amazon-eks","title":"Delhivery achieves 160 ms latency for high-precision geocoding using Amazon EKS","description":"Delhivery, a logistics provider in India, implemented a fine-tuned open-source Llama 3.2 1B large language model on Amazon EKS to support high-volume geocoding of pickup and drop-off addresses. The system processes up to 8,000 requests per minute at 160 milliseconds latency using NVIDIA A10G GPU-backed G5 Xlarge instances and the vLLM framework. Delhivery cut model-serving costs by approximately 80 percent and accelerated prototyping cycles from two days to under six hours, working with the AWS Prototyping and Cloud Engineering (PACE) team.","company":"Delhivery","industry":"Logistics","country":"India","aiCapabilities":["Large Language Models"],"technology":["Llama 3.2 1B","Amazon Elastic Kubernetes Service","Amazon EC2 G5 Instances","NVIDIA A10G GPU","NVIDIA Triton Inference Server","vLLM"],"deployment":"Public Cloud","problemStatement":"Delhivery initially tested serverless LLMs from third-party providers for high-precision geocoding of pickup and drop-off addresses, but these came with rate caps of 2,000 requests per minute or higher costs for provisioned access exceeding actual usage needs. Traditional machine learning models also lacked contextual understanding and required long training cycles, slowing experimentation and making it difficult to scale during demand spikes.","solutionApproach":"Delhivery selected and externally fine-tuned the open-source Llama 3.2 1B model, then engaged the AWS Prototyping and Cloud Engineering (PACE) team to identify suitable instance types, optimize model serving with NVIDIA Triton Inference Server, and package the deployment for Amazon EKS. Production deployment uses Amazon EKS with G5 Xlarge instances equipped with NVIDIA A10G GPUs and the vLLM framework, with auto scaling to handle demand.","businessValue":"The system processes up to 8,000 requests per minute at a latency of 160 milliseconds (measured at a concurrency of 30). Delhivery reduced its monthly model-serving costs by approximately 80 percent and accelerated prototyping cycles from two days to under six hours.","evidence":{"band":"high"},"sourceUrl":"https://aws.amazon.com/solutions/case-studies/delhivery-case-study/","dates":{"publishedAt":"2026-08-15T09:02:32.486Z","publishedAtSource":"ledger","updatedAt":"2026-08-26T09:41:52.459Z"},"license":"Open for reading and citing with a link to https://findausecase.com/use-cases/delhivery-achieves-160-ms-latency-for-high-precision-geocoding-using-amazon-eks. Bulk republication requires permission."}