Evals Are Broken, Use Them Anyway — Ara Khan, Cline
Evals are broken and often misleading, but engineers should still use them well in agentic workflows.
“evals are broken and you should use them anyway”
In a talk at AI Engineer, Cline's Ara Khan critiques two failure modes in AI evaluation: the 'objective metrics' camp that treats benchmark numbers as truth (benchmark maxing, where similar scores hide real model differences) and the 'taste' camp that over-relies on intuition. The argument is that despite being broken, evals remain valuable if engineers learn to interpret and leverage them within agentic flows. It's a practical engineering-culture signal about how the field measures model quality.