35 tracked signals on evals.
SimulationMaxxing: How Nubank ships agents 20× faster with simulations — Shreya Rajpal, Snowglobe
AI Engineer · Jul 29, 2026
Nubank ships AI customer support agents 20x faster by generating eval data in simulation rather than waiting for production data.
“If you generate your eval data in sim instead of waiting on production data you can ship agents 20x faster and we'll give you evidence for that.”
The Self-Driving Eval Trick No AI Benchmark Beats
LangChain · Sep 21, 2026
Manually reviewing 100-1,000 real examples beats any automated benchmark for evaluating AI models.
“there's just no better eval than looking at 100 examples or 1,000 examples”
When to Build Your Own Agent Harness | Harrison Chase, LangChain
Harrison Chase · Sequoia Capital · Aug 13, 2026
LangChain CEO argues companies must own their agent harness to truly own their AI intelligence
“The main job of a harness is to bring context to the model at the right point in time.”
SWE-bench is saturated, so Cognition built FrontierCode
LangChain · Aug 05, 2026
Cognition built FrontierCode eval after SWE-bench saturated, focusing on code mergeability
“okay, this code is technically correct. But would you actually merge it? Would you feel happy? Would this improve the quality of your code base?”
The Future of Evals: From LLM as a Judge to Agent as a Judge — Aparna Dhinakaran, Arize AI
AI Engineer · Jul 24, 2026
AI evals must evolve from LLM-as-judge to agent-as-judge as systems shift to multi-agent long-horizon tasks
“we didn't just make the problem harder, we actually got a fundamentally different type of problem”
How Salesforce Standardizes Agent Evals with LangSmith
LangChain · Jul 23, 2026
Salesforce AgentForce uses LangSmith to standardize evals across teams at scale
“LangSmith helps us to amplify our internal expertise, allowing us to run thousands and thousands of test cases.”
Production Evals For Agentic AI Systems - Nishant Gupta, Meta Superintelligence Labs
AI Engineer · Jun 25, 2026
Agentic AI evaluation must shift from model benchmarking to production-grade system behavior testing.
“The question is no longer did the model generate the right answer? The question is did the system behave correctly?”
Make Legal Write Your Evals: Building Jade, Chime’s Financial Copilot | Interrupt 26
LangChain · Jun 22, 2026
Chime built compliance evals for its AI co-pilot Jade by having legal teams co-author risk definitions.
“I would argue that evals are your alignment surface.”
How Lyft Builds Evals That Actually Matter in Production | Interrupt 26
LangChain · Jun 15, 2026
Lyft scaled to seven+ production AI agents at 35% resolution by building an offline-eval quality gate before shipping.
“You don't want to use your users as test data.”
The Return of the Data Scientist | Interrupt 26
LangChain · Jun 12, 2026
AI engineering evals are fundamentally data science, and teams keep repeating avoidable eval mistakes.
“I'm saying that using LLM judges kind of blindly without validating your validators is bad.”
Why Building an Eval Platform Is Harder Than It Looks — Braintrust
AI Engineer · Oct 06, 2026
Agent quality demands both pre-production evals and production observability to manage LLM non-determinism risks
“grading matters because LLMs are inherently non-deterministic”
From Vibes to Production: Evaluating and Shipping AI Agents That Work 101 — Laurie Voss, Arize AI
AI Engineer · Oct 05, 2026
Arize AI presents a 101 framework for evaluating AI agents from tracing to LLM-as-judge meta-evaluation
Evals in AI: A Deep Dive — Tejas Kumar, IBM
AI Engineer · Oct 05, 2026
AI evals provide reliability in advance, not just post-hoc validation
“Evals provide reliability, but not just any reliability, but reliability in advance.”
Agentic Search vs Vector Search for Coding Agents: We Ran the Eval — Braintrust
AI Engineer · Oct 05, 2026
Systematic evals beat gut-feel releases: 94% pass rates beat 'it looked good'
“I ran 200 different test scenarios, 94% of them passed, so we're releasing this feature.”
Inside Clay's Eval Stack: 300M Agent Runs, One LangSmith Pipeline
LangChain · Aug 28, 2026
Clay runs 300M agent executions monthly and calls evals non-negotiable at that scale
“eval's became non-negotiable”
AI Evals for Cross-Functional Teams — Nachiket Paranjape & Swaroop Chitlur Haridas, DoorDash
AI Engineer · Aug 28, 2026
DoorDash evolved AI evals from an engineering task into a cross-functional platform for non-engineers.
“cost is a number one concern these days”
Your Agent Evolved. Your Evals Didn't. — Ameya Bhatawdekar, Braintrust
AI Engineer · Aug 20, 2026
AI agents evolve rapidly but evaluation frameworks fail to keep pace with model changes
“building a demo with AI is really easy but making it production quality is really hard”
Why Do AI Agents Hallucinate?
LangChain · Jul 31, 2026
Agent loops amplify LLM hallucinations by compounding errors across downstream steps
“one hallucinated fact in step two can poison step three, step four and everything downstream”
Building Closed-Loop Evals for a Multimodal Agent at Scale — Soumya Gupta & Jai Chopra, Uber
AI Engineer · Jul 24, 2026
Uber built a multimodal AI agent to enhance merchant food photos at scale without looking AI-generated
“we're threading the needle here. We need to be able to stay faithful to the original image, preserve the brand of the merchant, and avoid everything looking the same.”
From Signal to PR: Anatomy of a Self-Improving Agent — Jason Lopatecki, Arize
AI Engineer · Jul 24, 2026
Arize founder argues observability is evolving from human dashboards to autonomous self-healing agent systems
“how do I build systems that autonomously fix themselves?”
Evaluating AI Agents: A production blueprint with Strands and AgentCore
AWS Machine Learning Blog · Jul 23, 2026
AWS and Motorway cut AI agent error rates from 1-in-8 to 1-in-50 with a production eval pipeline.
“The agent gives a confident-sounding response, but how do you prove it works reliably with real money on the line?”
Build Evals That Actually Matter - Nick Ung, Lyft
AI Engineer · Jul 19, 2026
Lyft applies ML model evaluation rigor to AI agents before production deployment.
“if we are running um offline evaluations for our machine learning model before that goes to productions, I think we should do the same for AI agents as well”
LangSmith: The Agent Engineering Platform
LangChain · Jul 17, 2026
LangSmith positions itself as a full-stack, agnostic agent engineering platform with autonomous improvement.
“LangSmith Engine is your agent for agent engineering.”
User Signal Dies at the Retrieval Boundary - Sonam Pankaj, StarlightSearch
AI Engineer · Jun 28, 2026
Agent failures stem from static retrieval that never learns from eval and observability signals.
“We made wrong answers appear faster and cheaper, but we forgot to make retrieval learn.”
Agents in Production: How OpenGov Built and Scaled OG Assist - Gabe De Mesa, OpenGov
AI Engineer · Jun 26, 2026
OpenGov built and scaled OG Assist, a production AI agent embedded across its government ERP product suite.
Evals Are Broken, Use Them Anyway — Ara Khan, Cline
AI Engineer · Jun 06, 2026
Evals are broken and often misleading, but engineers should still use them well in agentic workflows.
“evals are broken and you should use them anyway”
How to build agents when the smartest AI isn't smart enough
LangChain · Jun 04, 2026
Benchling's AI agents pair LLMs with structured lab data to roughly halve drug-discovery-to-patient time.
“What we've seen is when you kind of put these models on top of the right data, the quality answers go way up.”
Understand and fix Agent Framework apps with observability and evals | DEM361
Microsoft Developer (Build) · Jun 04, 2026
Most developers building agents lack visibility into what their agent does under the hood; observability and evals close that gap.
“you probably don't know what your agent is actually doing under the hood”
Build, Test, Deploy, Monitor: The Agent Development Lifecycle Explained
LangChain · Sep 30, 2026
LangChain defines a four-phase agent development lifecycle: build, test, deploy, monitor.
“What is the agent development lifecycle? So the agent development lifecycle is the process that agent engineers use to build and improve their agents.”
smevals - a small eval suite for evaluating models, prompts, and harnesses
Simon Willison · Jul 31, 2026
smevals is a new open-source eval framework for comparing AI model and prompt configurations
“I've been trying to figure out an approach I like for evals for several years now. smevals is my third iteration on the idea and it feels right to me.”
Evaling Video Slop — Maor Bril, Character.ai
AI Engineer · Jul 25, 2026
Video generation eval methods lag far behind generation quality, missing narrative coherence entirely.
“the hard part was never how to make video. The hard part was how do we generate good enough video and how do we judge if the video is good enough?”
How Evals and Prompts Shape Agent Behavior — Preetika Bhateja & Daniel Bump, YouTube Ads
AI Engineer · Jul 24, 2026
YouTube Ads team built agent reliability via layered evals, critique agents, and strong tool foundations.
“the reliability of your agent is basically a function of the capabilities of the agent uh the guard rails and the evals”
Agents Building Agents - Alfonso Graziano, Nearform
AI Engineer · Jun 28, 2026
NearForm uses AI agents and spec-driven development to iteratively build more reliable, secure AI agents.
“an agent is an LLM inside a an agentic loop, which is connected to tools and can retrieve context. That's it.”
Advanced Workshop: Mastering AI Observability — Doug Guthrie, Braintrust
AI Engineer · Oct 06, 2026
Braintrust workshop teaches isolating signal from AI agent trace noise using observability tooling
Are AI labs pelicanmaxxing?
Simon Willison · Jul 22, 2026
Systematic study finds no evidence AI labs train models to draw pelicans on bicycles better
“Pelicans aren't drawn any better than other animals. Bicycles aren't drawn any better than other vehicles. And no lab draws the combination better than its pelicans and bicycles already predict.”