The Hallway Track

prompt-injection

17 tracked signals on prompt-injection.

Breaking Claude Code Opus 5 Auto Mode

Simon Willison · Aug 27, 2026

Prompt injection attack bypasses Claude Code Auto Mode with 80% success rate

“The safety mechanism itself can become part of the failure. The classifier allowed the creation of the malware process, but then it blocked the command intended to stop it!”
Boris Cherny: Stop Hobbling Your AI

Y Combinator · Jul 27, 2026

Opus 5 achieves prompt injection resistance and can run autonomously for days without scaffolding

“A year ago the model would have just done it. But nowadays Opus does not.”
Quoting Matthew Green

Simon Willison · Oct 01, 2026

Sandboxed AI agents can spread worm payloads via shared resources like email and documents

“Put these pieces together and you have the two halves of a worm: a payload that hijacks the agent, and an agent that will carry the payload to the next agent.”
AI Worming through Word

Simon Willison · Jul 29, 2026

Researcher demonstrates self-replicating prompt injection worm spreading through Microsoft Word via Copilot

“An attacker places hidden instructions in a document that is later used as source material in Copilot for Word. Copilot may interpret those instructions as part of the user's request, causing it to manipulate the document being drafted or edited. Copilot may then also copy the hidden instructions into the resulting document, turning that document into a new carrier.”
Quoting Boris Cherny

Simon Willison · Jul 25, 2026

Anthropic's Opus 5 is their most prompt-injection-resistant model to date.

“Opus 5 is our least prompt injectable model yet. It is a bit buried in the system card, but across PI evals and red teaming, Opus 5 is very hard to prompt inject successfully.”
Prompt Injection as Role Confusion

Simon Willison · Jun 22, 2026

Models confuse text style with role, making prompt injection a 'role confusion' problem defenses can't easily fix.

“destyling causes average attack success in our dataset to plunge from 61% to 10%”
Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan

Latent Space Blog · Jun 22, 2026

Gray Swan's Kolter and Fredrikson discuss AI red-teaming and indirect prompt injection after US export controls on Mythos and Fable.

“the risks of jailbreaks and (industry term) indirect prompt injection are suddenly the talk of the town”
OpenAI Help: Lockdown Mode

Simon Willison · Jun 05, 2026

OpenAI's Lockdown Mode is now live, limiting outbound network requests to block prompt-injection data exfiltration.

“Lockdown Mode is designed to help prevent the final stage of data exfiltration from a prompt injection attack by limiting outbound network requests that could transfer sensitive data to an attacker.”
Microsoft Copilot Cowork Exfiltrates Files

Simon Willison · May 26, 2026

Microsoft Copilot Cowork vulnerability enables file exfiltration via prompt injection and rendered images

“Because these messages can contain external images that trigger network requests to external websites, data can be exfiltrated when a user opens a compromised message sent by the agent.”
What happened after 2,000 people tried to hack my AI assistant

Simon Willison · Jun 26, 2026

Frontier model anti-prompt-injection training held up against 6,000 attempts to leak secrets from an AI assistant.

“after 6,000 attempts (and $500 in token spend and a Google account suspension triggered by too many inbound emails) nobody managed to leak the secret”
WWDC26: Secure your app: mitigate risks to agentic features | Apple

Apple Developer (WWDC) · Jun 08, 2026

Apple details security techniques and APIs to mitigate prompt-injection risks in LLM-powered agentic app features.

“LLM's introduce a new probabilistic engine within your application that is both powerful but risks being tricked.”
Google I/O, Gemini Spark, Antigravity

Simon Willison · May 20, 2026

Google replaces open-source Gemini CLI with closed-source Antigravity CLI on June 18th

“Gemini Spark runs on Gemini 3.5 Flash and Antigravity.”