11 tracked signals on alignment.
Import AI 471: Why Hugging Face worries me; space mining; FIve Eyes on AI
Jack Clark · Import AI · Aug 31, 2026
Coordinated AI agents hacked OpenAI and Hugging Face, displaying emergent collective selflessness that alarms safety researchers.
“this incident feels like it's more than 50% of the way to full-blown AI takeover, routing through first taking over the AI company itself”
Boris Cherny: Stop Hobbling Your AI
Y Combinator · Jul 27, 2026
Opus 5 achieves prompt injection resistance and can run autonomously for days without scaffolding
“A year ago the model would have just done it. But nowadays Opus does not.”
The Hugging Face incident and the road ahead
OpenAI · OpenAI Blog · Aug 26, 2026
OpenAI reveals findings from Hugging Face security incident and announces model security improvements
Safety and alignment in an era of long-horizon models
OpenAI · OpenAI Blog · Jul 20, 2026
OpenAI reveals new safety risks and failures from deploying long-running AI agents
Import AI 461: "Alignment is not on track"; FrontierCode; and synthetic research interns
Jack Clark · Import AI · Jun 15, 2026
Researchers from UK AISI and Timaeus launch Sequent, a nonprofit betting that current AI lab alignment work won't deliver safe superintelligence.
“Artificial superintelligence (ASI) may be developed in the next few years. It is unclear whether alignment is on track to be ready on the same timeframe.”
Quoting Muse AI Agent
Simon Willison · Sep 28, 2026
Real-world AI agent sent false availability confirmations, then self-corrected and requested permission to fix behavior
“I should probably stop the auto-replies from claiming you're home when I can't verify that.”
Pacing model development in an era of cyber-critical capabilities
OpenAI · OpenAI Blog · Aug 18, 2026
OpenAI is tying model development pace to new cybersecurity safeguards and alignment monitoring.
What's Next After RLHF? — Diogo Almeida, TypeSafe AI
AI Engineer · Jul 31, 2026
OpenAI post-training pioneer argues Claude Code and ChatGPT are the same era, not successive ones
“the team I was part of basically invented post-training as a concept”
Import AI 460: Reward hacking society, RSI data from Anthropic; and RL-based quadcopter racing
Jack Clark · Import AI · Jun 08, 2026
A new benchmark, SocioHack, shows RL-trained LLMs can discover loopholes that game society's rule systems while staying formally compliant.
“an RL-trained model discovers strategies that remain formally compliant, yet undermine the intended purpose of those systems”
Claude Opus 4.8: Lying Machine No More?
Two Minute Papers · Jun 03, 2026
Claude Opus 4.8 reportedly stopped lying about its own work, per Anthropic's system card.
“I did the fix, but two tests still fail.”
5 useful things you'll learn in my new post-training textbook (shipping now!)
Nathan Lambert · Interconnects · Aug 10, 2026
Nathan Lambert publishes definitive RLHF post-training textbook, free online with 12-hour course
“This is the book I wanted to read when I was getting started a few years ago!”