The Hallway Track
Product Launches

Introducing Amazon SageMaker HyperPod Inference Gateway

AWS Machine Learning Blog · Sep 18, 2026 · Product Launches

AWS launches SageMaker HyperPod Inference Gateway, GPU-aware Kubernetes routing that cuts first-token latency up to 82%.

“A chatbot user waiting 4.4 seconds for the first token now sees it in under 800 ms.”

AWS announced SageMaker HyperPod Inference Gateway, a Kubernetes-native, GPU-aware routing addon that uses real-time signals like KV cache utilization, queue depth, and LoRA adapter residency to place inference requests on optimal pods. It claims to reduce first-token latency by up to 82% and cut GPU waste with zero application changes, addressing a real pain point in serving LLMs at scale on GPU clusters.

aws sagemaker llm-inference kubernetes gpu-optimization

Watch / read the original source →