KV Cache-Aware Routing and P/D Disaggregation on Kubernetes — Yuchen Fama & Ashish Kamra, Red Hat
Red Hat demonstrates KV cache-aware routing and prefill/decode disaggregation for agentic LLM workloads on Kubernetes.
“benchmarks actually don't show you is the chaotic reality of multi-turn interactions, massive context fluctuations which are very typical of agentic workloads”
Red Hat engineers present production challenges of running agentic LLM workloads, arguing that standard benchmarks obscure the real complexity of multi-turn, variable-context inference. They detail KV cache-aware routing and P/D disaggregation as key strategies, using GLM 5.2 as a case study on Kubernetes. Signals Red Hat's growing seriousness as an open-source inference infrastructure player (top vLLM contributor), but the talk is practitioner-level rather than a major announcement.