The Hallway Track
Research Findings

ITBench-AA: Frontier Models Score Below 50% on the First Benchmark for Agentic Enterprise IT Tasks — by Artificial Analysis and IBM

Hugging Face · Hugging Face Blog · May 27, 2026 · Research Findings

Frontier AI models score below 50% on new agentic enterprise IT benchmark ITBench-AA

IBM and Artificial Analysis released ITBench-AA, the first benchmark specifically targeting agentic AI performance on enterprise IT tasks, revealing that even frontier models fail more than half the time. This exposes a significant capability gap between current AI hype and real-world enterprise automation readiness. The finding is a meaningful signal for enterprise AI adoption timelines and sets a public evaluation standard for agentic IT workflows.

benchmarks agentic-ai enterprise-it ibm evaluation

Watch / read the original source →