The Hallway Track

infrastructure

100 tracked signals on infrastructure.

Quoting SpaceX S-1

Simon Willison · May 20, 2026

Anthropic signed $1.25B/month compute deal with SpaceX's Colossus clusters through May 2029

“Cloud Services Agreements with Anthropic PBC...with respect to access to compute capacity across COLOSSUS and COLOSSUS II...the customer has agreed to pay us $1.25 billion per month through May 2029”
The Industrialization of Intelligence

NVIDIA GTC · Sep 02, 2026

A new industrial revolution is unfolding powered by AI technologies.

“No company can build this infrastructure alone.”
[AINews] Memory prices up 500% in 12 months

Latent Space Blog · Aug 19, 2026

Memory prices up 500% in 12 months as hyperscalers lock in all 2027 DRAM production capacity

“Some are calling it the RAMpocalypse; I prefer "RAMageddon."”
Technical deep dive: AgentCore payments and innovation in agentic commerce

AWS Machine Learning Blog · May 26, 2026

AWS launches AgentCore Payments enabling AI agents to autonomously execute microtransactions via stablecoins

“Amazon Bedrock AgentCore payments is the first managed service within Amazon Bedrock AgentCore that helps AI agents autonomously execute microtransaction payments for paid APIs, MCPs, and content with a few lines of code.”
AI, Infrastructure, and the Next Investment Cycle

a16z · Sep 30, 2026

Hyperscalers are investing all short-term operating cash flow into AI capacity as demand outpaces supply

“Hyperscalers are investing all of their short-term operating cash flow into building this capacity to meet demand that continues to outpace supply in almost every case we see”
Teaching agents to pay — Anna Spysz, Stripe

AI Engineer · Sep 01, 2026

Stripe engineer demos autonomous shopping agent as agentic commerce infrastructure matures

“the infrastructure for agentic transactions has been laid down by companies like Google, OpenAI, and Stripe”
Taking Reinforcement Learning Cross Datacenter — Nan Jiang, Modal

AI Engineer · Aug 10, 2026

Modal proposes decoupling RL rollout workers from trainer clusters to use distributed GPU capacity across datacenters.

“IO wants all four of these at the same time. Enough GPU, same region, fast fabric, and available now. Any of these like is manageable, but all four of them that are pretty hard to get at the same time.”
Compute at Sea

Y Combinator · Jul 23, 2026

Y Combinator is seeking founders to build offshore AI compute flotillas on the ocean

“It sounds crazy, but we think part of the answer may be to move compute offshore.”
What Makes Open Models Fast in Production — Sujee Maniyam, Nebius

AI Engineer · Oct 03, 2026

Nebius, backed by $2B Nvidia investment, offers full-stack open LLM inference from silicon to service

“Typically, most AI teams are stuck between choosing two bad options. Closed APIs are very easy to get started with, but you often hit a ceiling very quickly.”
Build Applications on NVIDIA BlueField Faster with NVIDIA DOCA Agent Skills

NVIDIA Developer Blog · Oct 01, 2026

NVIDIA releases domain-specific DOCA Agent Skills for BlueField infrastructure development

“AI agents are becoming a standard part of development workflows, but general-purpose agents weren't built with specialized infrastructure software such as NVIDIA DOCA in mind.”
NVIDIA BlueField-4 Powers New Scale-In Network Infrastructure for Agentic AI Factories

NVIDIA Developer Blog · Aug 24, 2026

NVIDIA BlueField-4 DPU enables dedicated networking for agentic AI factory infrastructure

“Agentic AI factories connect diverse users, agents, applications, data sources, and storage systems to massively accelerated compute at multi-terabit bandwidth per server, making dedicated DPU processing essential for line-rate networking, storage, and security.”
How Generative Recommenders Are Redefining RecSys at Scale

NVIDIA Developer Blog · Aug 20, 2026

LLMs are shifting recommender systems from embedding-similarity to generative next-action prediction

“The advent of LLMs has inspired a shift from the traditional embedding-similarity-based objective to a generative one, where the goal is to predict the next action or item in a large catalog given a sequence of user histories.”
ModelExpress: Distributing Model Artifacts at the Speed of Light

NVIDIA Developer Blog · Jul 24, 2026

NVIDIA ModelExpress accelerates multi-hundred-GB model weight distribution across GPU clusters

“Every byte moved has a cost. As model checkpoints grow to hundreds of gigabytes or even a terabyte, that cost adds up quickly.”
Optimize model training on Amazon SageMaker AI with NVIDIA Blackwell

AWS Machine Learning Blog · Jun 25, 2026

AWS SageMaker AI now offers P6-B200 instances with NVIDIA Blackwell GPUs to optimize large model training.

“models that previously required multi-node setups can run on a single 8-GPU node, which means faster iteration cycles, less networking overhead, and lower infrastructure costs”
What Is a Sandbox? (And Why Every Agent Needs One) | Spill The Tea

LangChain · Jun 17, 2026

Sandboxes are isolated virtual computers that let agents run arbitrary code safely and scale via parallelization.

“Sandboxes are isolated virtual computers which your agents can use to write code, run commands, or automate browser tasks.”
AgentWatch: Proactive AWS monitoring with ambient agents

AWS Machine Learning Blog · May 26, 2026

AWS introduces AgentWatch, an ambient AI agent for proactive infrastructure monitoring via Slack

“The agent works continuously alongside your team to observe your infrastructure, analyze patterns, and surface insights without requiring constant human intervention.”
Scale AI with Google's TPU software stack

Google Developers (Google I/O) · May 21, 2026

Google reveals TPU V8 splits into training (8T) and inference (8I) specialized variants

“a lot of the intelligence is actually coming from inference”
Railway: The Agent-Native Cloud — Jake Cooper

Latent Space Blog · May 20, 2026

Railway is repositioning as agent-native cloud infrastructure with 70% margins and 3M users.

“the activation energy to ship something to production should be near zero.”