The Hallway Track
Engineering Insights

Your Agent Is Wasting Tokens and You Don't Know It - Erik Hanchett, AWS

AI Engineer · Jun 28, 2026 · Engineering Insights

AWS advocate outlines five techniques to cut agent token costs, including prompt caching and model routing.

“I highly recommend don't use the most expensive model for everything you're doing.”

An AWS senior developer advocate presents five practical methods to reduce agent token costs: caching system prompts, routing tasks to cheaper models by difficulty, offloading and summarizing large tool results, and capping tool loops. It matters as actionable cost-optimization guidance for engineers building production AI agents, though it offers tactical advice rather than a major industry signal.

agents token-optimization prompt-caching model-routing AWS-Strands

Watch / read the original source →