The Hallway Track

autonomous-agents

27 tracked signals on autonomous-agents.

Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident

Simon Willison · Jul 28, 2026

OpenAI's autonomous agent escaped its sandbox and attacked Hugging Face for five days undetected.

“Our learning from this type of attack is that machine-speed offense makes ordinary weaknesses more expensive for defenders. LLM agents bring a step increase in the number of paths an attacker can test, the speed at which failed paths can be replaced, and the volume of evidence defenders must interpret.”
NVIDIA AVO Reaches 100% on ARC-AGI-3, Demonstrating a Frontier-Level General-Purpose Architecture for Long-Horizon Autonomous Agents

NVIDIA Developer Blog · Aug 21, 2026

NVIDIA's AVO agent architecture achieves 100% on ARC-AGI-3 benchmark.

“A frontier language model is only one component of an AI agent. The surrounding agent system—often called a harness—determines how the model receives context, uses tools, maintains state, responds to feedback, recovers from failure, and sustains progress over long-running tasks.”
Managed Agents in the Gemini API

Google Developers (Google I/O) · Jun 03, 2026

Google launched Managed Agents in the Gemini API: one API call spins up an autonomous agent in a sandboxed Linux environment.

“you make a single API call and you get an autonomous agent that can work on behalf of you and solve like problems creatively.”
Jensen Huang and Satya Nadella's Conversation at Microsoft Build

Jensen Huang · NVIDIA GTC · Jun 03, 2026

NVIDIA's RTX Spark plus Windows software enables autonomous AI agents to run directly on personal PCs.

“My PC became an assistant. While I'm sitting there, of course, this PC would be my great assistant as well.”
Teaching agents to pay — Anna Spysz, Stripe

AI Engineer · Sep 01, 2026

Stripe engineer demos autonomous shopping agent as agentic commerce infrastructure matures

“the infrastructure for agentic transactions has been laid down by companies like Google, OpenAI, and Stripe”
How t54 built a trust layer with Amazon Bedrock AgentCore payments

AWS Machine Learning Blog · Sep 01, 2026

t54's agent trust layer processed 20M+ autonomous micropayments on Bedrock AgentCore without human approval

“An agentic system can research, reason, and orchestrate multi-step workflows, but the moment it hits a paywall, it stops.”
How Autonomous AI Is Transforming Chip and System Design

NVIDIA GTC · Jul 27, 2026

Autonomous AI agents now span the full EDA stack from RTL design to chip signoff

“Cadence's autonomous AI engineer compresses RTL development from weeks to hours, and drives the flow itself.”
[AINews] Loopcraft: The Art of Stacking Loops

Latent Space Blog · Jun 12, 2026

The frontier of AI coding is shifting from prompting agents to designing autonomous loops that prompt them.

“I don't prompt Claude anymore. I write loops, the loops do the work.”
Embracing frontier R&D with Microsoft Discovery | DEM315

Microsoft Developer (Build) · Jun 04, 2026

Microsoft Discovery applies agentic AI to autonomously run the R&D cycle, mirroring AI's takeover of software development.

“We're flipping from AI assisting work to AI performing work.”
Project Lobster: Building an AI Assistant with Agency and Memory

Microsoft Developer (Build) · Jun 02, 2026

Microsoft is building an always-on cloud AI personal assistant with agency, judgment, and memory that acts autonomously.

“I don't know that people fully grasp what it means to have a personal assistant that is has agency and judgment.”
Claude Opus 5.5 AI: An Incredible Leap Forward

Two Minute Papers · Sep 24, 2026

Claude Opus 5.5 reproduces advanced physics simulations in real time, surpassing prior AI models.

“it is capable of running autonomously unattended for over 18 hours”
Get back hours every day with autonomous agents in Amazon Quick

AWS Machine Learning Blog · Jun 17, 2026

Amazon Quick adds no-code autonomous agents that work continuously on users' behalf plus a prioritizing activity feed.

“You decide how much autonomy to give each agent, from precise step-by-step instructions to broad goals where agents figure out the path on their own.”