Agent reliability is a control problem, not a prompting problem — harness the LLM, don't let it drive.
“Reliability was never a prompting problem. It's a control problem.”
27 tracked signals on cost-optimization.
Agent reliability is a control problem, not a prompting problem — harness the LLM, don't let it drive.
“Reliability was never a prompting problem. It's a control problem.”
Clay runs over 350 million go-to-market AI agents monthly, processing trillions of tokens per week.
“We run this over 350 million times a month. It processes trillions of tokens every week.”
LangChain achieved 64% cost reduction in their coding agent via model routing with no quality loss
“we were able to see a 64% reduction in median cost per thread with no measurable change in quality”
GPT-6 improves prompt caching with higher hit rates, explicit breakpoints, and new latency/cost controls.
Databricks cut $1M annually in AI agent waste within a single hour of optimization work.
Fable's high cost ends the era of relying on new models to cheaply solve engineering problems
“Prior to Fable, it felt silly to waste too much time improving your coding harness or context strategies. A new model would arrive at the same price (or cheaper!) and paper over most of your problems.”
Model routing is now a critical AI deployment strategy driven by frontier cost and open-weights competition
“A big goal of Glean is to avoid using LLMs for tasks where we don't need them. Sometimes you'll see queries in Glean where people are adding two numbers or multiplying two numbers. They could have used a calculator to do that.”
OpenAI GPT-5.6 models launch on Amazon Bedrock with explicit prompt caching at 90% discount.
“Cached input is billed at a 90 percent discount (see the Amazon Bedrock pricing page) and stays available for reuse for 30 minutes.”
Microsoft Foundry distills production agent traces into smaller fine-tuned models that run cheaper and faster.
“the question is not just whether my agent works. It is about whether I can afford to run my agent a 100 million times”
Microsoft Foundry model router achieves 37.7% cost savings by routing prompts to appropriately-sized models
Databricks Unity AI Gateway offers smart routing to cut costs 30%+ while matching frontier model quality
Agentic AI systems should route tasks across frontier and open models for cost efficiency.
“being able to build these agents that are able to use a system of models to address the problem at hand and do it most cost-effectively is going to benefit us”
Notion is repositioning as a human-agent collaboration platform managing token costs sustainably.
“Today that collaboration happens between humans and agents. Humans and humans, agents and agents.”
Pairing Amazon Nova 2 Lite with Claude Sonnet 4.6 in a two-model Bedrock pipeline cuts per-page document-digitization cost by about two-thirds.
“This two-model approach costs about two-thirds less per page than a single-model alternative that sends the entire task to one vision-language model.”
Developers building cost-efficient AI apps over-optimize the model instead of the end-to-end system and data access.
“developers get right into thinking about optimizing the model versus optimizing the system”
Post-training open-source reasoning models with reinforcement learning keeps agent costs flat while quality improves in Foundry.
“deploying a model into production is just the beginning”
Fine-tuned small open-source reasoning models on Foundry can match frontier models for your domain at a fraction of the cost.
“This is almost like test-driven development for the age of agents.”
A new AI called Jev makes decisions instead of generating text, running up to 200x faster than chatbots.
“It pops out instantly, up to about 200 times faster than our chatbots today.”
AWS proposes using a cheaper filter model to compress RAG context before the primary model call
Async invocation patterns eliminate idle compute costs when calling Bedrock AgentCore agents from serverless pipelines.
Developers waste effort optimizing the model instead of the end-to-end system and data access.
“when AI has access to the right data, it spends more time reasoning and coming up with the right outcomes and higher quality outcomes”
Running AI SRE agents on all enterprise alerts naively costs ~$730,000/year, driven mostly by input-token costs.
“these are all driven by the cost on input tokens”
Qualcomm argues small on-device tasks like docstrings shouldn't be routed to large cloud models.
“if you want to predict the future, you have to go and create it”
Databricks Lakebase offers cost optimization strategies for managed Postgres workloads
AWS webinar previews new Bedrock cost allocation features and CUR columns for tracking generative AI spend.
“We are excited to talk about some new features that have come out. How to look at your spend using the cur and these new columns”
Azure free tiers and serverless offerings enable viable apps under $25/month
“you're not just optimizing your architecture, but you're really building an economical, sustainable and scalable solution”
Azure free tiers and serverless offerings enable viable AI apps under $25/month
“You really don't want to sort of trouble your model with all kinds of questions and expensive queries.”