The Hallway Track

evals

35 tracked signals on evals.

The Self-Driving Eval Trick No AI Benchmark Beats

LangChain · Sep 21, 2026

Manually reviewing 100-1,000 real examples beats any automated benchmark for evaluating AI models.

“there's just no better eval than looking at 100 examples or 1,000 examples”
SWE-bench is saturated, so Cognition built FrontierCode

LangChain · Aug 05, 2026

Cognition built FrontierCode eval after SWE-bench saturated, focusing on code mergeability

“okay, this code is technically correct. But would you actually merge it? Would you feel happy? Would this improve the quality of your code base?”
How Salesforce Standardizes Agent Evals with LangSmith

LangChain · Jul 23, 2026

Salesforce AgentForce uses LangSmith to standardize evals across teams at scale

“LangSmith helps us to amplify our internal expertise, allowing us to run thousands and thousands of test cases.”
The Return of the Data Scientist | Interrupt 26

LangChain · Jun 12, 2026

AI engineering evals are fundamentally data science, and teams keep repeating avoidable eval mistakes.

“I'm saying that using LLM judges kind of blindly without validating your validators is bad.”
Evals in AI: A Deep Dive — Tejas Kumar, IBM

AI Engineer · Oct 05, 2026

AI evals provide reliability in advance, not just post-hoc validation

“Evals provide reliability, but not just any reliability, but reliability in advance.”
Why Do AI Agents Hallucinate?

LangChain · Jul 31, 2026

Agent loops amplify LLM hallucinations by compounding errors across downstream steps

“one hallucinated fact in step two can poison step three, step four and everything downstream”
Build Evals That Actually Matter - Nick Ung, Lyft

AI Engineer · Jul 19, 2026

Lyft applies ML model evaluation rigor to AI agents before production deployment.

“if we are running um offline evaluations for our machine learning model before that goes to productions, I think we should do the same for AI agents as well”
LangSmith: The Agent Engineering Platform

LangChain · Jul 17, 2026

LangSmith positions itself as a full-stack, agnostic agent engineering platform with autonomous improvement.

“LangSmith Engine is your agent for agent engineering.”
How to build agents when the smartest AI isn't smart enough

LangChain · Jun 04, 2026

Benchling's AI agents pair LLMs with structured lab data to roughly halve drug-discovery-to-patient time.

“What we've seen is when you kind of put these models on top of the right data, the quality answers go way up.”
Evaling Video Slop — Maor Bril, Character.ai

AI Engineer · Jul 25, 2026

Video generation eval methods lag far behind generation quality, missing narrative coherence entirely.

“the hard part was never how to make video. The hard part was how do we generate good enough video and how do we judge if the video is good enough?”
Agents Building Agents - Alfonso Graziano, Nearform

AI Engineer · Jun 28, 2026

NearForm uses AI agents and spec-driven development to iteratively build more reliable, secure AI agents.

“an agent is an LLM inside a an agentic loop, which is connected to tools and can retrieve context. That's it.”
Are AI labs pelicanmaxxing?

Simon Willison · Jul 22, 2026

Systematic study finds no evidence AI labs train models to draw pelicans on bicycles better

“Pelicans aren't drawn any better than other animals. Bicycles aren't drawn any better than other vehicles. And no lab draws the combination better than its pelicans and bicycles already predict.”