The Hallway Track

agent-evaluation

14 tracked signals on agent-evaluation.

Introducing: LangSmith Tuned Evaluators

LangChain · Aug 18, 2026

LangChain launches tuned evaluators that beat frontier models on agent evals at lower cost

“our post-trained perceived error evaluator outperformed all frontier closed and open models, while still remaining the most cost-effective”
Evaluate AI agents systematically with Agent-EvalKit

AWS Machine Learning Blog · Jun 11, 2026

AWS released Agent-EvalKit, an open-source toolkit for systematically evaluating AI agents' full execution paths.

“An agent might deliver a well-structured, actionable response while hallucinating, fabricating facts because its tools returned empty results.”
How Do You Actually Evaluate an AI Agent?

LangChain · Aug 26, 2026

Agent evals must run continuously using LLM-as-judge since non-deterministic outputs make exact-match testing impossible.

“The teams that really get this right treat eval sort of as a muscle. You build it early, you're running it constantly because I think the alternative is like finding out your agent tried recommending the wrong product three weeks ago from a customer complaint.”
Evaluating Deep Agents using LangSmith on AWS

AWS Machine Learning Blog · May 28, 2026

LangSmith on AWS provides five evaluation patterns for deep agents in production

“Validating AI agent behavior before production is one of the hardest problems in applied AI.”
What is LangSmith?

LangChain · Sep 22, 2026

LangSmith is LangChain's framework-agnostic platform for tracing, testing, deploying, and monitoring LLM agents.

“Agents are a black box, making them difficult to observe and debug.”