The Future of Evals: From LLM as a Judge to Agent as a Judge — Aparna Dhinakaran, Arize AI
AI evals must evolve from LLM-as-judge to agent-as-judge as systems shift to multi-agent long-horizon tasks
“we didn't just make the problem harder, we actually got a fundamentally different type of problem”
Aparna Dhinakaran of Arize AI argues that first-gen eval frameworks built for single-prompt LLMs are now mismatched to the complex multi-agent, long-horizon systems teams are shipping in 2024-2025. Drawing on 100M+ evals/month across customer deployments, she highlights that online evals on production traces—not just offline benchmarks—are the critical feedback loop for catching emergent failure modes. The shift from 'LLM as a Judge' to 'Agent as a Judge' reflects a fundamental change in what's being evaluated, not just added complexity.