The Hallway Track
Engineering Insights

Build Evals That Actually Matter - Nick Ung, Lyft

AI Engineer · Jul 19, 2026 · Engineering Insights

Lyft applies ML model evaluation rigor to AI agents before production deployment.

“if we are running um offline evaluations for our machine learning model before that goes to productions, I think we should do the same for AI agents as well”

Lyft engineers Nick Ung and Aka share their evaluation framework for a customer support AI agent system built over 1-2 years, covering offline evals, online evals, and an eval harness. Their core thesis is that AI agents deserve the same rigorous pre-production evaluation discipline already applied to traditional ML models. This is a practitioner-level engineering talk with actionable signal for teams operationalizing agentic systems.

evals AI agents customer support MLOps production AI

Watch / read the original source →