The Hallway Track
Engineering Insights

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

OpenAI · OpenAI Blog · Jul 29, 2026 · Engineering Insights

Two API settings tripled OpenAI's GPT-5.6 scores on the ARC-AGI-3 benchmark

OpenAI revealed that enabling two specific API settings—retaining reasoning and enabling compaction—tripled GPT-5.6's performance on the ARC-AGI-3 benchmark. This is significant because ARC-AGI-3 is a leading test of general reasoning and fluid intelligence, making this a notable efficiency and capability leap. The finding suggests meaningful headroom in existing models that practitioners may be leaving on the table.

ARC-AGI-3 GPT-5.6 benchmark reasoning OpenAI efficiency

Watch / read the original source →