From Vibes to Production: Evaluating and Shipping AI Agents That Work 101 — Laurie Voss, Arize AI
Arize AI presents a 101 framework for evaluating AI agents from tracing to LLM-as-judge meta-evaluation
Laurie Voss, head of developer relations at Arize AI and co-founder of NPM, delivered a hands-on seminar covering the full agent evaluation pipeline: tracing raw data, running deterministic and LLM-based evals, building custom LLM-as-judge assessors, and validating judges via meta-evaluation. The session used Claude Agent SDK to build a demo agent and Claude Code for automated response improvement from eval feedback. It matters because eval tooling for agents remains a key unsolved pain point, and this provides a concrete, reproducible workflow practitioners can adopt today.