LangChain introduces eval capabilities for Managed Deep Agents.
evaluation
37 tracked signals on evaluation.
LangSmith CLI enables evaluation of user frustration with LLMs.
“It'll then produce some feedback, which gets assigned on the trace, and consists of a score plus a reasoning.”
Google DeepMind is piloting the world's first double-blind AI evaluations
Traditional single-number benchmarks fail to capture that modern AI capability scales with test-time compute budget.
“The capability of the model is a function of how much money you put into it.”
Frontier AI models score below 50% on new agentic enterprise IT benchmark ITBench-AA
NVIDIA's SWE-Serve benchmark tests whether AI coding agents' patches work in live model serving, not just local tests.
“An AI coding agent’s patch can pass tests yet fail when the server loads a real model and handles requests.”
OpenAI introduces MentalHealthBench, an expert-informed benchmark for evaluating safe AI responses in mental health conversations.
“MentalHealthBench is an expert-informed benchmark for evaluating helpful and safe AI responses across realistic mental health conversations.”
Strands Evals and Amazon Bedrock AgentCore add skill-focused evaluators to measure agent skill selection and instruction following.
“A skill is a reusable set of instructions, usually stored in a SKILL.md file, that teaches an agent a domain-specific task like redacting a contract, reconciling an invoice, or following a team’s pull-request conventions.”
Ufonia's Dora AI has completed 200,000 clinical calls and is contracted to reach 1 million patients.
“the model card won't save you. Um you can't claim like some model vendors said that they have 92% on some benchmark. Um it's not a defense at a post incident review.”
Standard computer use benchmarks are gameable by blind replay scripts, invalidating frontier model comparisons
“if you try to evaluate this kind of agent on standard benchmarks such as OSWorld or MobileWorld, you will see that the success rate of this agent compared to the frontier model from which the agent was extracted is actually the same or even better”
AI labs optimize for benchmark scores rather than real-world model quality, warns industry insider
“unfortunately the teams are not getting better models overall but better Elm Marina models whatever that is possibly something with a lot of nested list bullet points and emojis”
Current AI safety policies fail to account for test-time compute, where model capability scales with money spent.
“The capability of the model is a function of how much money you put into it, basically.”
OpenAI launches LifeSciBench, an expert-authored benchmark for evaluating AI on real-world life science research tasks.
“Introducing LifeSciBench, an expert-authored, expert-reviewed benchmark for evaluating how AI systems handle real-world life science research tasks and decisions.”
Cognition launches FrontierCode, a benchmark grading code quality and maintainability over passing-test 'slop'.
“Many SWE-bench-Passing PRs Would Not Be Merged into Main”
SWE-rebench evaluates coding agents monthly on fresh, decontaminated real-world software engineering tasks.
“if you want to build some open truly decontaminated benchmark, um time splits are the only way”
Microsoft Foundry's agent tracing and evaluations hit GA and now extend to any agent framework and deployment target.
“Shipping an AI agent is the easy part. Keeping it accurate, safe, and accountable in production is where teams get stuck.”
Voice agent transcripts mask real failures that only audio-level observability reveals
“the logs lie”
AWS launches managed evaluation framework for multi-agent systems with explainability as a first-class dimension
“Traditional evaluation approaches that focus only on model response quality are insufficient for agentic systems, where correctness depends on tool selection, workflow execution, and adherence to business constraints.”
AI agent evaluation must shift from scoring single tool calls to measuring full multi-step task completion.
“Scoring whether the model sounds right tells you almost nothing about whether the work finished.”
Microsoft Foundry offers end-to-end tooling for deploying, governing, and monitoring AI agents in production.
“The biggest challenge when building your AI agent is how do you get it in production?”
NVIDIA launches SkillEvaluator to benchmark whether agent skills actually improve performance
“AI agents are only as effective as the context they receive.”
Current LLM benchmarks fail to measure learning ability across sequential tasks over time.
“imagine that every time you do something, you completely forget your memory... That's the premise under which we're evaluating language models today.”
Taste Labs exits stealth to build data infrastructure ending AI slop in subjective domains
“our whole mission is basically how do we end AI slop?”
DeepSWE is a contamination-resistant coding benchmark with 113 original tasks across ~100 repositories
AI self-improvement loops require domain-specific evaluation functions, not just agentic design.
“does the code compile or not”
Building production AI systems requires a four-phase framework—requirements, system design, evaluation, optimization—not just vibe coding.
“Specs are the new code. The art is in defining the product requirements, the system design, and evaluate criteria so you can be confident that your AI coding buddies are building the right thing.”
Hugging Face explores benchmarking open models for agentic capability against your own tooling.
AWS's Strands Evals SDK adds detectors that automatically diagnose AI agent failures and recommend fixes.
“Detectors answer “why did it fail?” by producing diagnoses at the per-span level with categorized failures, causal chains, and fix recommendations.”
Hugging Face introduces olmo-eval, an evaluation workbench for the model development loop.
AWS released the open source Nova Sonic Test Harness to automatically evaluate voice agents at scale without a microphone.
“It runs complete multi-turn conversations with Amazon Nova Sonic automatically, evaluates them using LLM-as-judge techniques, and can even detect cases where the model's audio output doesn't match its text output (audio hallucinations).”
Etsy built a production gifting agent on LangChain with a thin ReAct harness that drove high purchase rates.
“What we found was that it returns high-quality search results with a relatively thin harness”
No good standards exist for agent development; real-world continual learning is largely absent.
“there aren't really good standards for these things, right? Like there there's not some one-size-fits-all uh solution”
Databricks releases OfficeQA Pro V2 benchmark for enterprise grounded-reasoning evaluation
Credit Genie debugs thousands of LangGraph agent traces using LangSmith's insights segmentation feature.
“One thing we like about working with the LangChain team is that it feels more similar to working with an internal team than an external company.”
Building reliable AI systems requires observability, evaluation, and experimentation, not magic.
“It's really the same set of patterns, just maybe a different flavor coming out. And it's really it feels like magic, but it's not magic, right? It's all just engineering.”
Microsoft Foundry's agent platform unifies tracing, evaluation, and optimization to test non-deterministic AI agents end-to-end.
“The properties that makes agents so useful like they're stateful, they they're long horizon, they can plan, they can course correct, they interact with their environment through tools are also the ones that makes it very hard to test.”
Hugging Face published research on benchmark optimization in speech recognition