Amazon Web Services

AWS Inferentia

AWS Inferentia chips are designed by AWS to deliver high performance at the lowest cost in Amazon EC2 for deep learning and generative AI inference applications. The first-generation AWS Inferentia chip powers Amazon EC2 Inf1 instances, which deliver up to 2.3x higher throughput and up to 70% lower cost per inference than comparable Amazon EC2 instances. AWS Inferentia2 delivers up to 4x higher throughput and up to 10x lower latency compared to Inferentia, and Inferentia2-based Amazon EC2 Inf2 instances are optimized to deploy large language models and latent diffusion models at scale, and are the first inference-optimized instances in Amazon EC2 to support scale-out distributed inference with ultra-high-speed connectivity between chips. The AWS Neuron SDK integrates natively with popular frameworks such as PyTorch and TensorFlow to deploy models on Inferentia chips (and train them on AWS Trainium chips) with minimal code changes.

Official product page