Import AI 469: Science AI; RSI simulator; and Zuck's technological pessimism
DiG-bench shows Fable 5 displays creative intuition; current frontier models cannot beat discovery games
“The new frontier for analyzing AI systems is understanding how good they are at inferring the unwritten rules of their environment”— Mark Zuckerberg
DiG-bench is a new 70-game benchmark from an Oxford/MIT/Princeton consortium testing AI systems' ability to infer hidden rules through exploration rather than instruction, with all games remaining unbeatable by today's frontier models. Fable 5 with Claude Code is singled out as showing early creative intuition on the benchmark, a notable callout for Anthropic's newest model. The benchmark uses private, handcrafted games to prevent training contamination, making it a more robust evaluation signal than typical public leaderboards.