Apple launches Core AI framework for on-device model inference across iOS and macOS.
AI Hardware & Edge AI Summit
- Dates
- 2026-09-09 → 2026-09-11
- Location
- San Jose, CA
- Ecosystem
- chip infra
- Importance
- 7/10
Related coverage & signals
Baseten raised $13B Series F as inference engineering emerges as a critical AI discipline
“How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?”
Nebius, backed by $2B Nvidia investment, offers full-stack open LLM inference from silicon to service
“Typically, most AI teams are stuck between choosing two bad options. Closed APIs are very easy to get started with, but you often hit a ceiling very quickly.”
vLLM open-source inference engine now runs on half a million GPUs simultaneously
“VLM is a inference engine. It is kind of like databases and operating system other critical software to power AGI.”
NVIDIA introduces the 'AI grid' reference design, positioning telcos to become intelligence providers in the token economy.
“There has truly never been a better time to be in Telecom.”
Google reveals TPU V8 splits into training (8T) and inference (8I) specialized variants
“a lot of the intelligence is actually coming from inference”
Databricks has built a proprietary inference platform serving frontier AI models at scale
“At Databricks, we've built a unique inference platform that serves every frontier”
Cerebras positions inference speed and cost as the critical battleground for AI infrastructure
“output is where speed and cost directly impact your users”
NVIDIA offers guidance on sizing GPUs for AI inference to optimize TCO
NVIDIA and OpenAI are building the largest data center ever at gigawatt scale.
“When it's fully built out, it will be the largest data center ever built.”
OpenAI's custom Jalapeño chip delivers industry-leading AI inference speed and efficiency.
AMD acquires Taalas, signaling Lisa Su's conviction in custom ASICs for AI inference
OpenAI and Broadcom unveiled Jalapeño, a custom AI chip built for LLM inference.
Anthropic signed $1.25B/month compute deal with SpaceX's Colossus clusters through May 2029
“Cloud Services Agreements with Anthropic PBC...with respect to access to compute capacity across COLOSSUS and COLOSSUS II...the customer has agreed to pay us $1.25 billion per month through May 2029”
California high-speed rail project estimated to cost $236 billion without completed segments.
“Governor Nuomo in a private video... doesn't believe that this project will ever be done in our lifetime.”
NVIDIA is shaping infrastructure for agentic AI with a new platform called Vera Rubin.
“The value gets generated in the computing and the transformation of the actual data into a calculation.”
Specialized GPU kernel generation can significantly enhance inference efficiency.
Meta introduces ZGateway, a proxy for ZippyDB traffic management.
A new industrial revolution is unfolding powered by AI technologies.
“No company can build this infrastructure alone.”
Memory prices up 500% in 12 months as hyperscalers lock in all 2027 DRAM production capacity
“Some are calling it the RAMpocalypse; I prefer "RAMageddon."”
OpenAI GPT-5.6 Sol, Terra, and Luna are now generally available on Amazon Bedrock.
NVIDIA Rubin GPU architecture is purpose-built for agentic AI workloads at scale
“What began as discrete AI model training and human-facing chat interfaces has evolved into always-on AI factories dedicated to producing intelligence at scale.”
Google's Gemma 4 brings frontier-level, multimodal, agentic AI to offline edge and mobile devices.
“The North Star of our smaller models is intelligence per byte of memory footprint.”
Databricks launched Omnigent, an open-source 'meta-harness' layer on top of the agentic stack to make agents effective at scale.
“we call it a meta harness of harnesses, if you know what an agent harness is.”
Clay runs over 350 million go-to-market AI agents monthly, processing trillions of tokens per week.
“We run this over 350 million times a month. It processes trillions of tokens every week.”
NVIDIA's Jensen Huang frames the AI factory as the largest infrastructure buildout in history and the best enterprise investment of the next decade.
“you can't think if you don't generate words”
Apple's new Core AI framework lets developers run advanced AI models entirely on-device in their apps.
“With Core AI, you can build app experiences where user's data never leaves their device.”
NVIDIA's RTX Spark plus Windows software enables autonomous AI agents to run directly on personal PCs.
“My PC became an assistant. While I'm sitting there, of course, this PC would be my great assistant as well.”
Fireworks and Baseten reach decacorn status as AI inference infrastructure sees explosive growth
“if you are gonna do multimodel inference, you are gonna need a router”
AWS launches AgentCore Payments enabling AI agents to autonomously execute microtransactions via stablecoins
“Amazon Bedrock AgentCore payments is the first managed service within Amazon Bedrock AgentCore that helps AI agents autonomously execute microtransaction payments for paid APIs, MCPs, and content with a few lines of code.”
Google releases EmbeddingGemma 2, a sub-billion open model unifying text, image, video, and audio embeddings on-device.
“a picture of a cat, the word cat, and the sound of a cat are all mapped close to each other in a shared high-dimensional embedding space”
AWS AgentCore Runtime Instances enable GPU-backed, 14-day multi-agent sessions on persistent EC2
“serverless sessions that cap at a few hours don't cut it”
Hyperscalers are investing all short-term operating cash flow into AI capacity as demand outpaces supply
“Hyperscalers are investing all of their short-term operating cash flow into building this capacity to meet demand that continues to outpace supply in almost every case we see”
Databricks launches Lakebase Search with full-text and vector search natively in Postgres for AI agents
NVIDIA releases open reference platform for continuous hardware-level AI agent safety monitoring
Google DeepMind adds private, server-side memory to Private AI Compute for personal AI.
“Introducing private, server-side memory to Private AI Compute for personal AI.”
A startup is betting on diffusion-based LLMs because they are inherently more parallel at inference time than autoregressive transformers.
“the bitter lesson is that the more parallel solution is the one that is eventually going to win”
OpenAI evolved its inference load balancer from engine-signal feedback loops to an explicit, predictable routing policy.
“our system evolved from routing based on feedback loops driven by engine signals to a more explicit and a predictable policy which is still informed by engine signals”
Moonshot AI's Kimi K3, a 2.8-trillion-parameter open-weight model, is now available on Amazon Bedrock.
“Kimi K3 is its most capable model and the first open model to reach 2.8 trillion parameters.”
Inference is no longer the bottleneck in AI development.
“Just do what I want.”