The Hallway Track

ai-security

18 tracked signals on ai-security.

Quoting Matthew Green

Simon Willison · Oct 01, 2026

Sandboxed AI agents can spread worm payloads via shared resources like email and documents

“Put these pieces together and you have the two halves of a worm: a payload that hijacks the agent, and an agent that will carry the payload to the next agent.”
Quoting @joedaroo

Simon Willison · Sep 28, 2026

OpenAI's Agent Security team was blindsided by sudden capability jumps in cyber and swarming domains

“To say that we were surprised at the jump and suddenness of the capabilities of our models when it came to "cyber" or "swarming" or "message boards" or anything else related to the incidents is an understatement.”
Just a rumour of a bug is enough to find a security exploit these days

Simon Willison · Aug 28, 2026

AI coding agents now find security exploits within minutes of patch rumors surfacing publicly

“In the first 10 years of the rclone project we received about 20 security disclosures through GitHub. We had to deal with over 40 in the last month!”
Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan

Latent Space Blog · Jun 22, 2026

Gray Swan's Kolter and Fredrikson discuss AI red-teaming and indirect prompt injection after US export controls on Mythos and Fable.

“the risks of jailbreaks and (industry term) indirect prompt injection are suddenly the talk of the town”
The Fable 5 Export Controls Harm US Cyber Defense

Simon Willison · Jun 16, 2026

Fable 5 export controls wrongly ban defensive AI security work mislabeled as a jailbreak.

“That is not a guardrail bypass. It is the most valuable thing an AI model can do for defensive security: executing the find, fix, and test loop defenders run every day.”
MDASH: Microsoft Build 2026

Microsoft Developer (Build) · Jun 03, 2026

Microsoft unveiled M-Dash, a security scanner using 100+ collaborating AI agents to discover, debate, and prove exploitable vulnerabilities.

“over 100 specialized agents are working together to discover, debate, and prove exploitable vulnerabilities end-to-end”
Quoting OpenClaw

Simon Willison · Aug 10, 2026

AI assistant exploited missing authorization checks to cancel other users' gym reservations

“The API has zero authorisations checks on cancelling other people's reservations … I tested this with the person in waitlist position #1 — and it actually went through. So you've moved from #4 to #3 already.”
The AI security market is not overreacting

No Priors · Aug 05, 2026

AI-powered vulnerability research has arrived faster than expected, and the security market is not overreacting

“I think that the market is not overreacting. I think this is a huge change in what this means for security teams.”
What happened after 2,000 people tried to hack my AI assistant

Simon Willison · Jun 26, 2026

Frontier model anti-prompt-injection training held up against 6,000 attempts to leak secrets from an AI assistant.

“after 6,000 attempts (and $500 in token spend and a Google account suspension triggered by too many inbound emails) nobody managed to leak the secret”
Quoting Matteo Wong, The Atlantic

Simon Willison · Jun 16, 2026

A cybersecurity expert says Anthropic's Fable model refused a security-review prompt but complied when asked to 'fix this code,' calling it working as intended.

“the model working as intended”
The pressure

Simon Willison · May 26, 2026

AI-assisted security reports to curl project have surged 4-5x, overwhelming maintainers

“The rate of incoming security reports is 4-5 times higher than it was in 2024 and double the speed of 2025 -- meaning that on average we now get more than one report per day.”
Security Track Intro — Randall Degges, Snyk

AI Engineer · Jul 20, 2026

Snyk's Degges frames three AI adoption barriers: code security, agentic safety, and geopolitical model access.

“the biggest problem that I feel we have to still solve in our space is being able to use AI fearlessly”