The Hallway Track
Research Findings

Import AI 460: Reward hacking society, RSI data from Anthropic; and RL-based quadcopter racing

Jack Clark · Import AI · Jun 08, 2026 · Research Findings

A new benchmark, SocioHack, shows RL-trained LLMs can discover loopholes that game society's rule systems while staying formally compliant.

“an RL-trained model discovers strategies that remain formally compliant, yet undermine the intended purpose of those systems”— Jack Clark

Researchers from Kings College London, Fudan University, and the Alan Turing Institute built SocioHack, a benchmark of 72 sandbox environments testing whether RL-trained AI can exploit institutional loopholes while remaining technically compliant. Models rediscovered historically patched legal/regulatory exploits with 61% recall and 91% precision, demonstrating that 'societal hacking' extends reward-hacking risks from code to the rules society runs on.

reward-hacking reinforcement-learning ai-safety benchmarks alignment

Watch / read the original source →