The Hallway Track

AI Hardware & Edge AI Summit

chips inference edge-ai infrastructure

Dates
2026-09-09 → 2026-09-11
Location
San Jose, CA
Ecosystem
chip infra
Importance
7/10

Official site →

Related coverage & signals

Introducing Core AI

Apple Developer (WWDC) · Aug 17, 2026

Apple launches Core AI framework for on-device model inference across iOS and macOS.

What Makes Open Models Fast in Production — Sujee Maniyam, Nebius

AI Engineer · Oct 03, 2026

Nebius, backed by $2B Nvidia investment, offers full-stack open LLM inference from silicon to service

“Typically, most AI teams are stuck between choosing two bad options. Closed APIs are very easy to get started with, but you often hit a ceiling very quickly.”
AI Grid 101: Top 5 Things You Need to Know

NVIDIA GTC · Jun 09, 2026

NVIDIA introduces the 'AI grid' reference design, positioning telcos to become intelligence providers in the token economy.

“There has truly never been a better time to be in Telecom.”
Scale AI with Google's TPU software stack

Google Developers (Google I/O) · May 21, 2026

Google reveals TPU V8 splits into training (8T) and inference (8I) specialized variants

“a lot of the intelligence is actually coming from inference”
Reliable LLM Inference at Scale

Databricks Blog · May 27, 2026

Databricks has built a proprietary inference platform serving frontier AI models at scale

“At Databricks, we've built a unique inference platform that serves every frontier”
Cerebras Explains | What Is AI Inference?

Cerebras · Sep 30, 2026

Cerebras positions inference speed and cost as the critical battleground for AI infrastructure

“output is where speed and cost directly impact your users”
[AINews] AMD buys Taalas

Lisa Su · Latent Space Blog · Aug 07, 2026

AMD acquires Taalas, signaling Lisa Su's conviction in custom ASICs for AI inference

Quoting SpaceX S-1

Simon Willison · May 20, 2026

Anthropic signed $1.25B/month compute deal with SpaceX's Colossus clusters through May 2029

“Cloud Services Agreements with Anthropic PBC...with respect to access to compute capacity across COLOSSUS and COLOSSUS II...the customer has agreed to pay us $1.25 billion per month through May 2029”
The Industrialization of Intelligence

NVIDIA GTC · Sep 02, 2026

A new industrial revolution is unfolding powered by AI technologies.

“No company can build this infrastructure alone.”
[AINews] Memory prices up 500% in 12 months

Latent Space Blog · Aug 19, 2026

Memory prices up 500% in 12 months as hyperscalers lock in all 2027 DRAM production capacity

“Some are calling it the RAMpocalypse; I prefer "RAMageddon."”
Inside NVIDIA Rubin GPU Architecture: Powering the Era of Agentic AI

NVIDIA Developer Blog · Jul 21, 2026

NVIDIA Rubin GPU architecture is purpose-built for agentic AI workloads at scale

“What began as discrete AI model training and human-facing chat interfaces has evolved into always-on AI factories dedicated to producing intelligence at scale.”
Gemma 4 in Action: Bringing Frontier AI to the Edge

Google Developers (Google I/O) · Jun 29, 2026

Google's Gemma 4 brings frontier-level, multimodal, agentic AI to offline edge and mobile devices.

“The North Star of our smaller models is intelligence per byte of memory footprint.”
Jensen Huang and Satya Nadella's Conversation at Microsoft Build

Jensen Huang · NVIDIA GTC · Jun 03, 2026

NVIDIA's RTX Spark plus Windows software enables autonomous AI agents to run directly on personal PCs.

“My PC became an assistant. While I'm sitting there, of course, this PC would be my great assistant as well.”
Technical deep dive: AgentCore payments and innovation in agentic commerce

AWS Machine Learning Blog · May 26, 2026

AWS launches AgentCore Payments enabling AI agents to autonomously execute microtransactions via stablecoins

“Amazon Bedrock AgentCore payments is the first managed service within Amazon Bedrock AgentCore that helps AI agents autonomously execute microtransaction payments for paid APIs, MCPs, and content with a few lines of code.”
Introducing EmbeddingGemma 2: An open model for natively multimodal embeddings

Google Developers (Google I/O) · Oct 06, 2026

Google releases EmbeddingGemma 2, a sub-billion open model unifying text, image, video, and audio embeddings on-device.

“a picture of a cat, the word cat, and the sound of a cat are all mapped close to each other in a shared high-dimensional embedding space”
AI, Infrastructure, and the Next Investment Cycle

a16z · Sep 30, 2026

Hyperscalers are investing all short-term operating cash flow into AI capacity as demand outpaces supply

“Hyperscalers are investing all of their short-term operating cash flow into building this capacity to meet demand that continues to outpace supply in almost every case we see”
Betting on Diffusion

No Priors · Sep 20, 2026

A startup is betting on diffusion-based LLMs because they are inherently more parallel at inference time than autoregressive transformers.

“the bitter lesson is that the more parallel solution is the one that is eventually going to win”
Introducing Kimi K3 on Amazon Bedrock

AWS Machine Learning Blog · Sep 18, 2026

Moonshot AI's Kimi K3, a 2.8-trillion-parameter open-weight model, is now available on Amazon Bedrock.

“Kimi K3 is its most capable model and the first open model to reach 2.8 trillion parameters.”