How Evals and Prompts Shape Agent Behavior — Preetika Bhateja & Daniel Bump, YouTube Ads
YouTube Ads team built agent reliability via layered evals, critique agents, and strong tool foundations.
“the reliability of your agent is basically a function of the capabilities of the agent uh the guard rails and the evals”
Engineers from YouTube Ads shared a practical framework for building reliable production agents: optimize individual LLM-friendly tools first, then add an independent critique agent with a remediation loop for self-correction, and finally establish robust evals to prove the value of incremental changes. The talk emphasizes that evals are essential for climbing the quality ladder in non-deterministic generative AI systems. While grounded and actionable, this is practitioner-level engineering guidance rather than a novel research or product announcement.