OpenAI's custom Jalapeño chip delivers industry-leading AI inference speed and efficiency.
NVIDIA GTC Spring 2026
- Dates
- 2026-03-16 → 2026-03-20
- Location
- San Jose, CA
- Ecosystem
- chip infra
- Importance
- 9/10
Related coverage & signals
NVIDIA Rubin GPU architecture is purpose-built for agentic AI workloads at scale
“What began as discrete AI model training and human-facing chat interfaces has evolved into always-on AI factories dedicated to producing intelligence at scale.”
NVIDIA Vera Rubin and Blackwell set new performance-per-watt benchmarks for agentic AI workloads
“across 100 trillion tokens of real-world usage, OpenRouter's State of AI report found that average prompt tokens per request grew roughly fourfold”
NVIDIA Groq 3 LPX accelerator enables ultrafast interactive inference on Vera Rubin NVL72 at long context
“NVIDIA Vera Rubin NVL72, the most versatile machine ever built, delivering high throughput and interactivity across the widest range of AI workloads—from small to large models, both open and closed.”
Fractile is building inference chips targeting 3-6 month architectural leads over competitors
“if you find a way to structurally secure a 3-6 month lead, you will win in all these implementations.”
Cerebras claims it has eliminated the speed-throughput tradeoff in AI inference
“In AI, speed is productivity.”
US AI supply chain faces extreme concentration risk beyond chips, especially in robotics components dominated by China
“Robotics is an incredibly promising industry and the supply chain is right now completely dominated by China.”
Google reveals TPU V8 splits into training (8T) and inference (8I) specialized variants
“a lot of the intelligence is actually coming from inference”
Cerebras runs models at over 1000 tokens per second for near-instantaneous inference
“We run models at over 1000 tokens per second, so thinking seems instantaneous.”
LLM inference is bottlenecked by memory bandwidth, not compute speed
“LLM inference is actually limited by memory bandwidth, not computational speed.”
AMD acquires World Labs for $8.2B, gaining spatial AI and Atlas sparse reconstruction model
“Recently we released Atlas, a first of its kind omni model architecture that solves a key outstanding problem in spatial intelligence: new camera view prediction.”
Sam Altman expects OpenAI to declare AGI achieved internally by December 2026.
“Automated AI Research Intern”
OpenAI's Jalapeño chip claims 1.5–1.9x better performance per watt than NVIDIA GB200/GB300.
AMD acquires Taalas, signaling Lisa Su's conviction in custom ASICs for AI inference
OpenAI and Broadcom unveiled Jalapeño, a custom AI chip built for LLM inference.
Mass Magnetics aims to recycle rare earth magnets for robotics and defense.
“A huge portion of the world's magnets will come from recycled materials.”
Specialized GPU kernel generation can significantly enhance inference efficiency.
Atlas is a new next-generation model capable of high-quality world simulation.
“Yesterday was a big day. You launched a new cutting-edge model that received a great reception.”
NVIDIA Vera Rubin delivers 10x measured performance per watt versus Blackwell on CoreWeave
“Not projections, not estimates, but real silicon measurements.”
Memory prices up 500% in 12 months as hyperscalers lock in all 2027 DRAM production capacity
“Some are calling it the RAMpocalypse; I prefer "RAMageddon."”
Apple launches Core AI framework for on-device model inference across iOS and macOS.
Google DeepMind launches Gemini Robotics 2 with whole-body intelligence, dexterity, and multi-robot collaboration.
“What we're building here is the intelligence layer to power any robot to do a broad range of useful tasks.”
NVIDIA declares physical AI has arrived after 15 years, unveiling full robotics stack and Japan mechatronics partnership.
“After 15 years of work, physical AI is here.”
Google DeepMind launches Gemini Robotics ER 2 with multi-robot collaboration and video understanding.
OpenAI GPT-5.6 Sol, Terra, and Luna are now generally available on Amazon Bedrock.
Black Forest Labs launches FLUX 3 Video, beating Seedance 2.0, Gemini Omni, and Grok Imagine
NVIDIA Cosmos is an open foundation model generating synthetic data for physical AI training.
“For physical AI, compute is data.”
Apple sued OpenAI for trade secret theft tied to its $6.5B hardware push
“OpenAI believes the product's defining feature will be its personality and ability to connect on a human-like level with users.”
NVIDIA's ENPIRE framework lets coding agents run a closed-loop self-improvement process for real-world robots, hitting 99% on dexterous tasks.
“Frontier coding agents can autonomously develop a policy to achieve a 99% success rate on challenging, dexterous manipulation tasks in the real world, such as PushT, organizing pins into a pin box, and using a cutter to cut a zip tie,”
Clay runs over 350 million go-to-market AI agents monthly, processing trillions of tokens per week.
“We run this over 350 million times a month. It processes trillions of tokens every week.”
NVIDIA's Jensen Huang frames the AI factory as the largest infrastructure buildout in history and the best enterprise investment of the next decade.
“you can't think if you don't generate words”
NVIDIA unveils Cosmos 3, an open omnimodel for physical AI that perceives, generates, and acts.
“Cosmos, the foundation for developers of the age of physical AI.”
NVIDIA releases Cosmos 3, billed as the first open omni-model for physical AI reasoning and action.
Fireworks and Baseten reach decacorn status as AI inference infrastructure sees explosive growth
“if you are gonna do multimodel inference, you are gonna need a router”
General-purpose LLM agents may outperform specialized models for robot control
“We are working on creating LLMs that drive robots”
Skydio uses agent orchestration to let one operator control large fleets of autonomous drones simultaneously.
“Thousands of these drones are now deployed across the country—at energy companies, public safety agencies, and construction firms.”
Dyna Robotics is building reliable, commercially deployable general-purpose manipulation policies, not just demos.
“Robot Demos Are Easy. Reliability Is Hard”
Robotics has appeared 'almost here' for 70 years despite current physical-AI hype reaching its peak.
“robotics has been "almost here" for the last 70 years.”
Perceptron AI wants to unify VLMs, VLAs, and world models into 'embodied foundational models' for real-time physical-world interaction.
“we want to move away from the distinction between VLM, VLA, world models, whatever you want to call it, to what we call embodied fundamental models”
Today's top vision models fail at basic visual reasoning, relying on pattern recognition over spatial understanding.
“there's actually a big gap between how these models handle visual thinking and how humans do it”