The Hallway Track
Engineering Insights

Tiered KV cache for large LLMs on Amazon SageMaker HyperPod with Curvine

AWS Machine Learning Blog · Aug 12, 2026 · Engineering Insights

AWS tiered KV cache on SageMaker HyperPod delivers 2.7x TTFT improvement and 100% cross-Pod cache hit rate

“With this architecture, workloads that previously required P5 instances can run on lower-cost G6e instances, reducing per-endpoint cost.”

AWS details a three-tier KV cache architecture (GPU, CPU, shared NVMe via Curvine) on SageMaker HyperPod that enables cross-replica cache sharing for vLLM deployments. Benchmarks show up to 2.7x time-to-first-token improvement and 100% cross-Pod cache hit rates, with cross-node L2 read latency of ~56ms. The practical implication is significant cost reduction: workloads previously requiring expensive P5 GPU instances can shift to lower-cost G6e hardware.

KV cache LLM inference SageMaker AWS vLLM infrastructure cost HyperPod

Watch / read the original source →