Breaking Claude Code Opus 5 Auto Mode
Prompt injection attack bypasses Claude Code Auto Mode with 80% success rate
“The safety mechanism itself can become part of the failure. The classifier allowed the creation of the malware process, but then it blocked the command intended to stop it!”
Security researcher Johann Rehberger demonstrated an 80% success rate attack against Claude Code's Auto Mode — Anthropic's recently-defaulted safety mechanism for coding agents — by exploiting Python import shadowing via a malicious zip archive. More alarmingly, Auto Mode blocked Claude's own cleanup commands after it detected the compromise, inverting the safety guarantee. The finding reinforces that sandboxed execution environments remain essential for unattended AI coding agents.