The Hallway Track
Engineering Insights

What Makes Open Models Fast in Production — Sujee Maniyam, Nebius

AI Engineer · Oct 03, 2026 · Engineering Insights

Nebius, backed by $2B Nvidia investment, offers full-stack open LLM inference from silicon to service

“Typically, most AI teams are stuck between choosing two bad options. Closed APIs are very easy to get started with, but you often hit a ceiling very quickly.”

Nebius Token Factory presented their approach to deploying and optimizing open LLMs in production, positioning vertical integration from bare-metal GPU infrastructure to model-serving APIs as a key differentiator. The company, a Nasdaq-listed European AI cloud with $2B in Nvidia backing, is among the first to run Blackwell Ultra hardware in production and is targeting 5GW of Nvidia capacity by 2030. The talk frames the open vs. closed API tradeoff as a core tension for AI teams, positioning self-hosted open models as the path past customization ceilings.

open-models inference infrastructure nvidia nebius production-llm

Watch / read the original source →