The Hallway Track
Governance & Policy

Quoting Matteo Wong, The Atlantic

Simon Willison · Jun 16, 2026 · Governance & Policy

A cybersecurity expert says Anthropic's Fable model refused a security-review prompt but complied when asked to 'fix this code,' calling it working as intended.

“the model working as intended”

Amid the White House's escalating conflict with Anthropic, a report alleged a 'Fable jailbreak,' but security expert Katie Moussouris reviewed it and concluded the model behaved as intended for cyberdefense, refusing a 'review for security issues' prompt while complying with a 'fix this code' request. This matters because it shows AI safety behavior being weaponized in a political fight over export controls and the AI race.

anthropic fable ai-security jailbreaking white-house

Watch / read the original source →