The Hallway Track

evaluation

37 tracked signals on evaluation.

How SWE-Serve Exposes the Gap Between Local Tests and Live Serving

NVIDIA Developer Blog · Sep 23, 2026

NVIDIA's SWE-Serve benchmark tests whether AI coding agents' patches work in live model serving, not just local tests.

“An AI coding agent’s patch can pass tests yet fail when the server loads a real model and handles requests.”
Introducing MentalHealthBench

OpenAI · OpenAI Blog · Sep 23, 2026

OpenAI introduces MentalHealthBench, an expert-informed benchmark for evaluating safe AI responses in mental health conversations.

“MentalHealthBench is an expert-informed benchmark for evaluating helpful and safe AI responses across realistic mental health conversations.”
Evaluate skill-equipped agents with Strands Evals and Amazon Bedrock AgentCore

AWS Machine Learning Blog · Sep 22, 2026

Strands Evals and Amazon Bedrock AgentCore add skill-focused evaluators to measure agent skill selection and instruction following.

“A skill is a reusable set of instructions, usually stored in a SKILL.md file, that teaches an agent a domain-specific task like redacting a contract, reconciling an invoice, or following a team’s pull-request conventions.”
Computer Use at the Edge of the Statistical Precipice — Pierluca D'Oro, Programma Labs

AI Engineer · Aug 14, 2026

Standard computer use benchmarks are gameable by blind replay scripts, invalidating frontier model comparisons

“if you try to evaluate this kind of agent on standard benchmarks such as OSWorld or MobileWorld, you will see that the success rate of this agent compared to the frontier model from which the agent was extracted is actually the same or even better”
When Will The Benchmaxxing Plague End? — Nick Heiner, Surge AI

Andrej Karpathy · AI Engineer · Aug 02, 2026

AI labs optimize for benchmark scores rather than real-world model quality, warns industry insider

“unfortunately the teams are not getting better models overall but better Elm Marina models whatever that is possibly something with a lot of nested list bullet points and emojis”
AI benchmark scores don’t tell you what you think they do

No Priors · Jun 29, 2026

Current AI safety policies fail to account for test-time compute, where model capability scales with money spent.

“The capability of the model is a function of how much money you put into it, basically.”
Introducing LifeSciBench

OpenAI · OpenAI Blog · Jun 17, 2026

OpenAI launches LifeSciBench, an expert-authored benchmark for evaluating AI on real-world life science research tasks.

“Introducing LifeSciBench, an expert-authored, expert-reviewed benchmark for evaluating how AI systems handle real-world life science research tasks and decisions.”
Build 2026: From observability to ROI for AI agents on any framework

Microsoft AI Foundry Blog · Jun 03, 2026

Microsoft Foundry's agent tracing and evaluations hit GA and now extend to any agent framework and deployment target.

“Shipping an AI agent is the easy part. Keeping it accurate, safe, and accountable in production is where teams get stuck.”
Evaluating multi-agent systems for explainability and helpfulness with Amazon Bedrock AgentCore

AWS Machine Learning Blog · Oct 05, 2026

AWS launches managed evaluation framework for multi-agent systems with explainability as a first-class dimension

“Traditional evaluation approaches that focus only on model response quality are insufficient for agentic systems, where correctness depends on tool selection, workflow execution, and adherence to business constraints.”
How to Evaluate AI Agents From Tool Calls to Task Completion

NVIDIA Developer Blog · Sep 21, 2026

AI agent evaluation must shift from scoring single tool calls to measuring full multi-step task completion.

“Scoring whether the model sounds right tells you almost nothing about whether the work finished.”
What does it really take to ship an AI agent?

Microsoft Developer (Build) · Aug 28, 2026

Microsoft Foundry offers end-to-end tooling for deploying, governing, and monitoring AI agents in production.

“The biggest challenge when building your AI agent is how do you get it in production?”
AI System Design: From Idea to Production - Apoorva Joshi, MongoDB

AI Engineer · Jun 28, 2026

Building production AI systems requires a four-phase framework—requirements, system design, evaluation, optimization—not just vibe coding.

“Specs are the new code. The art is in defining the product requirements, the system design, and evaluate criteria so you can be confident that your AI coding buddies are building the right thing.”
AI Agent Failure Detection and Root Cause Analysis with Strands Evals

AWS Machine Learning Blog · Jun 15, 2026

AWS's Strands Evals SDK adds detectors that automatically diagnose AI agent failures and recommend fixes.

“Detectors answer “why did it fail?” by producing diagnoses at the per-span level with categorized failures, causal chains, and fix recommendations.”
Evaluate your Amazon Nova Sonic voice agent at scale, no microphone required

AWS Machine Learning Blog · Jun 08, 2026

AWS released the open source Nova Sonic Test Harness to automatically evaluate voice agents at scale without a microphone.

“It runs complete multi-turn conversations with Amazon Nova Sonic automatically, evaluates them using LLM-as-judge techniques, and can even detect cases where the model's audio output doesn't match its text output (audio hallucinations).”
How Credit Genie Debugs Thousands of Agent Traces with LangSmith

LangChain · Jul 27, 2026

Credit Genie debugs thousands of LangGraph agent traces using LangSmith's insights segmentation feature.

“One thing we like about working with the LangChain team is that it feels more similar to working with an internal team than an external company.”
LLM Observability, Evaluation, Experimentation Platform — Dat Ngo, Arize

AI Engineer · Jun 07, 2026

Building reliable AI systems requires observability, evaluation, and experimentation, not magic.

“It's really the same set of patterns, just maybe a different flavor coming out. And it's really it feels like magic, but it's not magic, right? It's all just engineering.”
Behind the Scenes: Accelerating the AI Agent DevOps Lifecycle with End-to-End | LIVE159

Microsoft Developer (Build) · Jun 05, 2026

Microsoft Foundry's agent platform unifies tracing, evaluation, and optimization to test non-deterministic AI agents end-to-end.

“The properties that makes agents so useful like they're stateful, they they're long horizon, they can plan, they can course correct, they interact with their environment through tools are also the ones that makes it very hard to test.”