The Hallway Track
Research Findings

Computer Use at the Edge of the Statistical Precipice — Pierluca D'Oro, Programma Labs

AI Engineer · Aug 14, 2026 · Research Findings

Standard computer use benchmarks are gameable by blind replay scripts, invalidating frontier model comparisons

“if you try to evaluate this kind of agent on standard benchmarks such as OSWorld or MobileWorld, you will see that the success rate of this agent compared to the frontier model from which the agent was extracted is actually the same or even better”

Pierluca D'Oro from Programmabs (formerly Meta Superintelligent Labs) demonstrates that current computer use benchmarks like OSWorld are fundamentally flawed: a simple replay script under 1MB can match or beat frontier models by exploiting benchmark determinism. The paper also shows that the widely-used pass@K metric is mathematically equivalent to evaluating such a replay agent, meaning it measures benchmark exploitability rather than genuine agent capability. This undermines trust in published computer use benchmark results industry-wide.

computer-use benchmarking evaluation agents Meta

Watch / read the original source →