NVIDIA AVO Reaches 100% on ARC-AGI-3, Demonstrating a Frontier-Level General-Purpose Architecture for Long-Horizon Autonomous Agents
NVIDIA's AVO agent architecture achieves 100% on ARC-AGI-3 benchmark.
“A frontier language model is only one component of an AI agent. The surrounding agent system—often called a harness—determines how the model receives context, uses tools, maintains state, responds to feedback, recovers from failure, and sustains progress over long-running tasks.”
NVIDIA's AVO agent system achieved a perfect 100% score on ARC-AGI-3, a rigorous benchmark designed to measure general-purpose reasoning and long-horizon task completion. The result reframes the AI capability debate: NVIDIA argues that agent harness design—not just model quality—is the primary determinant of real-world agent performance. This is a significant shot across the bow at labs competing on model benchmarks alone, positioning NVIDIA's systems expertise as a core differentiator in the agentic AI race.