The Hallway Track

benchmarking

17 tracked signals on benchmarking.

The first known runaway AI agent - or a very bad marketing stunt?

Simon Willison · Jul 23, 2026

OpenAI's AI agent escaped its sandbox and attacked Hugging Face during large-scale benchmark testing.

“Hugging Face has an enormous attack surface. They have more interfaces than I can count which run untrusted models and code.”
We Let Claude Code and Codex Race Human Researchers — Elie Bakouch, Prime Intellect

AI Engineer · Sep 26, 2026

Prime Intellect benchmarks Claude Code and Codex against human researchers on AI research tasks

“we don't have any benchmark to quantify whether that's true or not, right? And even more so, we don't have an independent benchmark from small laboratories to understand whether to expect this in the near future.”
Computer Use at the Edge of the Statistical Precipice — Pierluca D'Oro, Programma Labs

AI Engineer · Aug 14, 2026

Standard computer use benchmarks are gameable by blind replay scripts, invalidating frontier model comparisons

“if you try to evaluate this kind of agent on standard benchmarks such as OSWorld or MobileWorld, you will see that the success rate of this agent compared to the frontier model from which the agent was extracted is actually the same or even better”
Benchmarking LLM Inference at Scale with AIPerf

NVIDIA Developer Blog · Sep 18, 2026

NVIDIA introduces AIPerf, a tool for benchmarking LLM inference performance at scale.

“All of these paths have the same problem: single-process performance limits, Python's GIL capping concurrency”