NVIDIA and OpenAI are building the largest data center ever at gigawatt scale.
“When it's fully built out, it will be the largest data center ever built.”
100 tracked signals on infrastructure.
NVIDIA and OpenAI are building the largest data center ever at gigawatt scale.
“When it's fully built out, it will be the largest data center ever built.”
Anthropic signed $1.25B/month compute deal with SpaceX's Colossus clusters through May 2029
“Cloud Services Agreements with Anthropic PBC...with respect to access to compute capacity across COLOSSUS and COLOSSUS II...the customer has agreed to pay us $1.25 billion per month through May 2029”
California high-speed rail project estimated to cost $236 billion without completed segments.
“Governor Nuomo in a private video... doesn't believe that this project will ever be done in our lifetime.”
NVIDIA is shaping infrastructure for agentic AI with a new platform called Vera Rubin.
“The value gets generated in the computing and the transformation of the actual data into a calculation.”
Meta introduces ZGateway, a proxy for ZippyDB traffic management.
A new industrial revolution is unfolding powered by AI technologies.
“No company can build this infrastructure alone.”
Memory prices up 500% in 12 months as hyperscalers lock in all 2027 DRAM production capacity
“Some are calling it the RAMpocalypse; I prefer "RAMageddon."”
Databricks launched Omnigent, an open-source 'meta-harness' layer on top of the agentic stack to make agents effective at scale.
“we call it a meta harness of harnesses, if you know what an agent harness is.”
NVIDIA's Jensen Huang frames the AI factory as the largest infrastructure buildout in history and the best enterprise investment of the next decade.
“you can't think if you don't generate words”
AWS launches AgentCore Payments enabling AI agents to autonomously execute microtransactions via stablecoins
“Amazon Bedrock AgentCore payments is the first managed service within Amazon Bedrock AgentCore that helps AI agents autonomously execute microtransaction payments for paid APIs, MCPs, and content with a few lines of code.”
AWS AgentCore Runtime Instances enable GPU-backed, 14-day multi-agent sessions on persistent EC2
“serverless sessions that cap at a few hours don't cut it”
Hyperscalers are investing all short-term operating cash flow into AI capacity as demand outpaces supply
“Hyperscalers are investing all of their short-term operating cash flow into building this capacity to meet demand that continues to outpace supply in almost every case we see”
Databricks launches Lakebase Search with full-text and vector search natively in Postgres for AI agents
NVIDIA releases open reference platform for continuous hardware-level AI agent safety monitoring
Google DeepMind adds private, server-side memory to Private AI Compute for personal AI.
“Introducing private, server-side memory to Private AI Compute for personal AI.”
Stripe engineer demos autonomous shopping agent as agentic commerce infrastructure matures
“the infrastructure for agentic transactions has been laid down by companies like Google, OpenAI, and Stripe”
NVIDIA NVLink Fusion enables NVHBM memory for custom XPU accelerators at hyperscale
Agents will use the web 1,000x more than humans, requiring reinvented search infrastructure
“We started parallel with the bet that agents would do it a thousandx more than humans ever have.”
OpenAI CFO frames full-stack chip-to-product compounding as path to cheaper, scalable intelligence
Meta open-sourced MetaRoCE, a new RDMA transport protocol built for million-GPU AI clusters on Ethernet.
“The fabric sees packets, but the NIC sees intent.”
Agent infrastructure is now commoditized by cloud platforms, making context the new competitive frontier
“They're all taxes one has to pay in order to get an agent out there to play the game.”
Modal proposes decoupling RL rollout workers from trainer clusters to use distributed GPU capacity across datacenters.
“IO wants all four of these at the same time. Enough GPU, same region, fast fabric, and available now. Any of these like is manageable, but all four of them that are pretty hard to get at the same time.”
Baseten raised $13B Series F as inference engineering emerges as a critical AI discipline
“How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?”
Meta doubled GEM ads model training efficiency to 20-25% MFU while scaling FLOPs 4x in 12 months
Y Combinator is seeking founders to build offshore AI compute flotillas on the ocean
“It sounds crazy, but we think part of the answer may be to move compute offshore.”
Most production ML security breaches stem from basic infrastructure mistakes, not exotic AI attacks
“almost everything that is breaking in the production ML security isn't some exotic AI attack. It's the same boring infrastructure mistakes that we supposedly fixed years ago.”
As AI agents move to production, the core challenge shifts from model intelligence to running probabilistic agents on deterministic infrastructure.
“These systems are fundamentally probabilistic. Infrastructure is not allowed to be.”
Nebius, backed by $2B Nvidia investment, offers full-stack open LLM inference from silicon to service
“Typically, most AI teams are stuck between choosing two bad options. Closed APIs are very easy to get started with, but you often hit a ceiling very quickly.”
NVIDIA releases domain-specific DOCA Agent Skills for BlueField infrastructure development
“AI agents are becoming a standard part of development workflows, but general-purpose agents weren't built with specialized infrastructure software such as NVIDIA DOCA in mind.”
NVIDIA VSS Blueprint 3.3 reduces cost of building production-scale visual AI agents
AWS achieves 40% throughput gain for MoE reinforcement learning using EKS, EFA, and DeepEP
AI agent payments lack controls, requiring new infrastructure beyond legacy payment systems
“the missing infrastructure layer for AI payments”
NVIDIA Dynamo's shadow engine recovery restores LLM inference capacity in seconds instead of minutes
NVIDIA Spectrum-X Ethernet reengineers networking for giga-scale AI GPU clusters
NVIDIA BlueField-4 DPU enables dedicated networking for agentic AI factory infrastructure
“Agentic AI factories connect diverse users, agents, applications, data sources, and storage systems to massively accelerated compute at multi-terabit bandwidth per server, making dedicated DPU processing essential for line-rate networking, storage, and security.”
Warp built a cloud agent platform after hitting limits of local laptop-based AI coding agents.
“we realized that we had kind of reached the limits of what we could do on our laptops, and we wanted agents to do work that was more long-running, that was adapted to different constraints”
LLMs are shifting recommender systems from embedding-similarity to generative next-action prediction
“The advent of LLMs has inspired a shift from the traditional embedding-similarity-based objective to a generative one, where the goal is to predict the next action or item in a large catalog given a sequence of user histories.”
Web agents face a model capabilities overhang as infrastructure, not models, is now the bottleneck
“there's a huge model capabilities overhang in this category specifically, and you all here can hopefully solve it”
vLLM open-source inference engine now runs on half a million GPUs simultaneously
“VLM is a inference engine. It is kind of like databases and operating system other critical software to power AGI.”
Personal AI codegen breaks traditional cloud infrastructure, requiring fundamentally new approaches
“personal AI codegen breaks traditional cloud infrastructure”
NVIDIA Vera delivers faster encryption, compression, and recovery for agentic AI storage workloads.
“Storage is an active part of every agentic AI workflow.”
Databricks rebrands Delta Sharing as Open Sharing under the Linux Foundation with expanded agent and model sharing
“previously with Delta Sharing, you could share a Delta table, and that was useful, but it's not enough these days”
Hugging Face now hosts 3 million public models, a 150x increase in a few years.
“More than 30% of Fortune 500 use hugging face as a part of AI workflows.”
NVIDIA ModelExpress accelerates multi-hundred-GB model weight distribution across GPU clusters
“Every byte moved has a cost. As model checkpoints grow to hundreds of gigabytes or even a terabyte, that cost adds up quickly.”
OpenAI launches Project Camellia AI data center in Effingham County, Georgia
Pinterest built an agentic LLM tool using MCP and ReAct to auto-diagnose Apache Spark job failures
“Our vision for a diagnostics agent was to ask it simply, 'Why did a job fail?' and get back a deep research document which provides evidence on the root cause of the failure.”
Human-focused API architecture fails at agent scale, requiring ground-up rethinking of auth systems.
“the human-focused architecture doesn't scale well for agents”
An agent's true identity is its append-only log, not the model or runtime executing it.
“The agent is its data. It's specifically the log.”
AWS SageMaker AI now offers P6-B200 instances with NVIDIA Blackwell GPUs to optimize large model training.
“models that previously required multi-node setups can run on a single 8-GPU node, which means faster iteration cycles, less networking overhead, and lower infrastructure costs”
Sandboxes are isolated virtual computers that let agents run arbitrary code safely and scale via parallelization.
“Sandboxes are isolated virtual computers which your agents can use to write code, run commands, or automate browser tasks.”
Google announced a $1.5 billion investment to expand its Alabama data center campus in 2026-2027.
Together AI details research scaling long-context training to 5 million token sequence lengths via context parallelism.
Microsoft's Arm-based Cobalt 200 VMs enter preview, offering ~50% better price-performance than Cobalt 100 for agentic AI workloads.
“we have about 22 million developers on ARM”
AWS introduces AgentWatch, an ambient AI agent for proactive infrastructure monitoring via Slack
“The agent works continuously alongside your team to observe your infrastructure, analyze patterns, and surface insights without requiring constant human intervention.”
Google reveals TPU V8 splits into training (8T) and inference (8I) specialized variants
“a lot of the intelligence is actually coming from inference”
Slurm topology-aware scheduling is required to unlock full GB200 NVL72 exascale performance
“NVIDIA GB200 NVL72 delivers exascale compute in a single rack, unlocking real-time trillion-parameter models.”
Railway is repositioning as agent-native cloud infrastructure with 70% margins and 3M users.
“the activation energy to ship something to production should be near zero.”
AI agent demand is causing 10-100x user growth that database infrastructure cannot keep pace with
“the number of daily active users has doubled, tripled, 10 or 100 times. This is largely driven by the demand for AI agents. The infrastructure simply cannot keep up with this development.”
NVIDIA releases AICR v1.0 for open, verifiable GPU cluster configuration on Kubernetes
NVIDIA NeMo Agent Toolkit now supports Amazon S3 Vectors as a custom persistent memory backend