Stripe acquires OpenRouter for $7B, validating model routing as high-margin AI infrastructure
“it was generating $100 million in annualized gross profit”
17 tracked signals on model-routing.
Stripe acquires OpenRouter for $7B, validating model routing as high-margin AI infrastructure
“it was generating $100 million in annualized gross profit”
Enterprise CIOs will soon be accountable for every AI token spent, role by role
“Every CIO is going to need to answer for every incremental token, where do we put it?”
Model routing is taking off, threatening OpenAI and Anthropic's premium-priced demand assumption.
“the era of picking one model is over”
LangChain achieved 64% cost reduction in their coding agent via model routing with no quality loss
“we were able to see a 64% reduction in median cost per thread with no measurable change in quality”
Stripe acquired OpenRouter, with both companies committed to enabling new company creation over consolidation
“No. We didn't think about it at all.”
Databricks Unity Gateway CLI adds intelligent model routing and budget governance for enterprise AI coding agents
“Unity Gateway's intelligent routing was immediately activated.”
Microsoft Foundry's model router auto-selects the best AI model across 27 options from multiple providers
“what if I told you you can have a model router that selects the best model for you?”
Databricks' Unity Gateway CLI lets enterprises govern and cost-control AI coding agents like Claude Code and Codex.
“From 0 to 60% I use Claude Code and the models are available there. 60 to 80% of the time I use Codex and a specific system model. And if I'm hitting 80-100% of my budget, I want to use a cheaper model like Kimmy.”
Amdocs and AWS deployed a multi-agent telco support system saving 61,000 hours of wait time at PLDT.
“I'm choosing the right model to the right task... using Claude for reasoning, but using the Amazon Nova for these fast queries. And this is reduce our tokens by 95%.”
Visual Studio now lets developers swap AI models from OpenAI, Anthropic, and Llama at project level.
“Model choice becomes a project level decision instead of a platform default.”
Fable's high cost ends the era of relying on new models to cheaply solve engineering problems
“Prior to Fable, it felt silly to waste too much time improving your coding harness or context strategies. A new model would arrive at the same price (or cheaper!) and paper over most of your problems.”
Model routing is now a critical AI deployment strategy driven by frontier cost and open-weights competition
“A big goal of Glean is to avoid using LLMs for tasks where we don't need them. Sometimes you'll see queries in Glean where people are adding two numbers or multiplying two numbers. They could have used a calculator to do that.”
Cast AI's Kimchi agent routes tasks to cheap open-source models to eliminate developer token rationing
“Our job is not to prohibit a developer from using an agent for coding. Our job is to make it so they can use it as much as they want, whenever they want, without any restrictions.”
Microsoft Foundry model router achieves 37.7% cost savings by routing prompts to appropriately-sized models
AWS advocate outlines five techniques to cut agent token costs, including prompt caching and model routing.
“I highly recommend don't use the most expensive model for everything you're doing.”
LangChain demos 'Jev,' a fast, cheap System 1 classification model for agent routing, guardrails, and evals.
“System 1 models are a class of AI models built to make fast structured decisions that software can use directly.”
Databricks Unity AI Gateway centralizes access control, model routing, and cost tracking for AI models and agents.
“You define access and costs once in a single control panel instead of configuring each tool separately.”