The Hallway Track
Engineering Insights

KV Cache-Aware Routing and P/D Disaggregation on Kubernetes — Yuchen Fama & Ashish Kamra, Red Hat

AI Engineer · Aug 27, 2026 · Engineering Insights

Red Hat demonstrates KV cache-aware routing and prefill/decode disaggregation for agentic LLM workloads on Kubernetes.

“benchmarks actually don't show you is the chaotic reality of multi-turn interactions, massive context fluctuations which are very typical of agentic workloads”

Red Hat engineers present production challenges of running agentic LLM workloads, arguing that standard benchmarks obscure the real complexity of multi-turn, variable-context inference. They detail KV cache-aware routing and P/D disaggregation as key strategies, using GLM 5.2 as a case study on Kubernetes. Signals Red Hat's growing seriousness as an open-source inference infrastructure player (top vLLM contributor), but the talk is practitioner-level rather than a major announcement.

KV cache prefill-decode disaggregation agentic inference vLLM Kubernetes Red Hat open source

Watch / read the original source →