The Hallway Track

benchmark

14 tracked signals on benchmark.

NVIDIA AVO Reaches 100% on ARC-AGI-3, Demonstrating a Frontier-Level General-Purpose Architecture for Long-Horizon Autonomous Agents

NVIDIA Developer Blog · Aug 21, 2026

NVIDIA's AVO agent architecture achieves 100% on ARC-AGI-3 benchmark.

“A frontier language model is only one component of an AI agent. The surrounding agent system—often called a harness—determines how the model receives context, uses tools, maintains state, responds to feedback, recovers from failure, and sustains progress over long-running tasks.”
GLM-5.3: How Chinese labs keep stride with the frontier

Nathan Lambert · Interconnects · Aug 14, 2026

Z.ai's GLM-5.3 surpasses Claude Fable 5 and GPT-5.6-Sol on select benchmarks

“On many benchmarks the model has surpassed Moonshot AI's Kimi K3 and on some it's surpassed Claude Fable 5 or GPT-5.6-Sol.”
Open-weight AI just hit 2.8 trillion parameters…

Fireship · Jul 22, 2026

Moonshot AI's Kimi K3 is a 2.8T-parameter open-weight model matching frontier closed models on coding benchmarks.

“it has OpenAI and Anthropic terrified because its Trust Me Bro benchmark performance is on par with and in some cases beating Claude Fable and GPT 5.6 Soul”
[AINews] Much ado about Open Weights

Satya Nadella · Latent Space Blog · Jul 28, 2026

Moonshot AI's Kimi K3 2.8T MoE beats Opus 4.8, claiming best open-weights model title

“This is more than a model drop; it is a fairly complete recipe for large-scale agentic post-training and serving.”
Introducing MentalHealthBench

OpenAI · OpenAI Blog · Sep 23, 2026

OpenAI introduces MentalHealthBench, an expert-informed benchmark for evaluating safe AI responses in mental health conversations.

“MentalHealthBench is an expert-informed benchmark for evaluating helpful and safe AI responses across realistic mental health conversations.”
GLM 5.3: Powerful AI Is Becoming Almost Free

Two Minute Papers · Sep 01, 2026

GLM 5.3 Flash delivers near-frontier AI performance for free via open weights.

“I think within this year, in a few months, Fable might be surpassed by free AI systems.”
Six Agent Harness Capabilities for Higher Model Performance

NVIDIA Developer Blog · Jul 27, 2026

Agent harness architecture drives double-digit benchmark swings independent of model choice

“Harness design alone can account for double-digit swings in benchmark results and significant differences in token cost”
Introducing LifeSciBench

OpenAI · OpenAI Blog · Jun 17, 2026

OpenAI launches LifeSciBench, an expert-authored benchmark for evaluating AI on real-world life science research tasks.

“Introducing LifeSciBench, an expert-authored, expert-reviewed benchmark for evaluating how AI systems handle real-world life science research tasks and decisions.”