Operating Distributed Inference Systems at Scale — Nishant Gupta & Naman Ahuja, Meta
Inference is now hyperscale infrastructure whose orchestration/control plane, not models, captures the value.
“The infence traffic already outpaces the largest microservices in the world and the rate of growth is fastest of any workload we have ever seen.”
Meta infra engineers argue AI inference is following the cloud's 2008 trajectory but compressed, with value shifting up-stack to an emerging orchestration/control plane handling routing, KV cache management, prefill/decode disaggregation, and multimodel multiplexing. They stress that agentic workloads scale as users × calls × tokens (thousands of calls per autonomous workflow), so capacity can no longer be planned like microservices. This matters because it reframes AI competitiveness around ecosystem orchestration rather than just models and kernels.