Productionizing LLM Gateways: Architecture, Tradeoffs and Hard Lessons — Kanish Manuja, Twilio
Traditional retry and circuit breaker patterns fail for LLMs; per-request provider fallback is required.
“If you have a single model provider, their ceiling is your ceiling. Their outage is your outage.”
Twilio principal engineer Kanesh Manuja explains why standard reliability patterns like exponential backoff and circuit breakers are insufficient for LLM infrastructure, arguing that slow and expensive LLM calls demand per-request multi-provider fallback instead. He frames LLM gateway design as a four-way tradeoff between availability, latency, guardrails, and cost that cannot be simultaneously maximized during degradation. This is practical, opinionated production guidance for teams running LLM workloads at scale.