NVIDIA

NVIDIA Triton Inference Server

NVIDIA Dynamo-Triton (formerly NVIDIA Triton Inference Server) is open-source software that enables deployment of AI models across major frameworks including TensorRT, PyTorch, ONNX, OpenVINO, Python and RAPIDS FIL, with dynamic batching and concurrent execution, supporting real-time, batched, ensemble and audio/video streaming workloads on NVIDIA GPUs, non-NVIDIA accelerators, x86 and ARM CPUs.

Official product page