The Hallway Track

open-weights

46 tracked signals on open-weights.

GLM-5.3: How Chinese labs keep stride with the frontier

Nathan Lambert · Interconnects · Aug 14, 2026

Z.ai's GLM-5.3 surpasses Claude Fable 5 and GPT-5.6-Sol on select benchmarks

“On many benchmarks the model has surpassed Moonshot AI's Kimi K3 and on some it's surpassed Claude Fable 5 or GPT-5.6-Sol.”
[AINews] Muse Glimmer and Spark: Open Weights return Personal Superintelligence promise

Mark Zuckerberg · Latent Space Blog · Aug 11, 2026

Zuckerberg repositions Meta as champion of personal AI against enterprise-focused labs with open weights model release

“Meta is the company primarily focused on building personal superintelligence for everyone. Most other labs are focused on building AI for companies, governments, or other institutions, so if those labs lead, then the balance of power will favor larger institutions over individuals.”
Open letters about AI development

Jensen Huang · Simon Willison · Aug 02, 2026

Competing open letters expose industry fracture over open-weight AI models and development pacing

“We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.”
Kimi K3: The open-weights escalation

Nathan Lambert · Interconnects · Jul 20, 2026

Kimi K3 is the strongest open-weights model ever, closing the US-China gap to 3-5 months

“the open-to-closed or American-to-Chinese model performance gap has been reduced from the debated 6-9 months to something shorter, say 3-5 months.”
DeepSeek Just Made Closed AI Look Ridiculous

Two Minute Papers · Aug 19, 2026

DeepSeek 4 Pro achieves near-frontier quality with MIT-licensed open weights, pressuring closed AI labs.

“DeepSeek has MIT licensed open weights. Anyone can run the exact same model at their own price.”
Introducing Muse Glimmer

Simon Willison · Aug 10, 2026

Meta releases Muse Glimmer, a 30B open-weights agentic model under Apache 2.0 license

“Muse Glimmer achieves strong success rates on full-task benchmarks including DeepSearch QA, MCP-Atlas, 𝛕-Bench and SWE-Bench, which measure its ability to work within scaffolds, write and debug code, and resolve multi-turn requests from start to finish.”
Deploying Kimi K3 on AWS

AWS Machine Learning Blog · Jul 30, 2026

Moonshot AI's Kimi K3 is the first open-weight model in the 3 trillion parameter class

“Kimi K3, a 2.8 trillion parameter Mixture of Experts (MoE) model that represents the first open-weight system to reach the 3 trillion parameter class”
Kimi K3 Just Broke The Economics Of AI

Two Minute Papers · Jul 29, 2026

Kimi K3's open-weights 2.8T parameter model matches frontier quality at significantly lower API cost

“even if you don't ever use it, it will be pushing token prices down”
[AINews] Much ado about Open Weights

Satya Nadella · Latent Space Blog · Jul 28, 2026

Moonshot AI's Kimi K3 2.8T MoE beats Opus 4.8, claiming best open-weights model title

“This is more than a model drop; it is a fairly complete recipe for large-scale agentic post-training and serving.”
Quoting Thomas Ptacek

Simon Willison · Jul 22, 2026

2025 open-weight models could already execute sandbox escapes and network hacks with a pentest harness

“I genuinely believe that if you took an open weights model from 2025 and built a pentest harness for it, it could do this kind of sandbox escape and scan/hack in most networks. This is only surprising because you assume OpenAI has sounder sandboxes.”
GLM-5.2 is the step change for open agents

Nathan Lambert · Interconnects · Jun 22, 2026

Z.ai's open-weight GLM-5.2 marks a step-change for open agentic models, rivaling top labs.

“minor version numbers can have AI models crossing meaningful user experience thresholds”
DiffusionGemma

Simon Willison · Jun 10, 2026

Google released DiffusionGemma, an open-weight Apache 2 diffusion-based text generation model.

“That research has returned in the best possible way: as a new open weight (Apache 2 licensed) Gemma model”
[AINews] not much happened today

Latent Space Blog · Jun 05, 2026

NVIDIA released Nemotron 3 Ultra, a fully open 550B MoE model with 1M context optimized for agentic workloads.

“up to 5x faster”
Introducing Mistral Large 4: Le chonk

Simon Willison · Oct 06, 2026

Mistral releases 1 trillion parameter Large 4 model, reclaiming competitive position

“it's great to see Mistral put out a model that's back to being maybe about 6 months behind the frontier”
Introducing Hy4 Preview

Simon Willison · Aug 29, 2026

Tencent releases Hy4, a 770B open-weight LLM with 1M token context window

This Small AI Will Change Everything

Two Minute Papers · Aug 24, 2026

Qwen 3.8B runs on consumer laptops while matching frontier model performance

“if we wait a bit, we might get frontier level systems running on our laptops”
Frontier Model Cost and Open-Weights Popularity is Driving Demand for Model Routing

Latent Space Blog · Aug 18, 2026

Model routing is now a critical AI deployment strategy driven by frontier cost and open-weights competition

“A big goal of Glean is to avoid using LLMs for tasks where we don't need them. Sometimes you'll see queries in Glean where people are adding two numbers or multiplying two numbers. They could have used a calculator to do that.”
DeepSeek V4 Pro 0813 (on OpenRouter)

Simon Willison · Aug 12, 2026

DeepSeek V4 Pro 0813 launches with 1.7T open weights and reasoning-level-dependent outputs.

“I've not noticed this kind of difference from any other model”
[AINews] not much happened today

Latent Space Blog · Aug 01, 2026

DeepSeek V4-Flash 0731 matches GPT-5.6 performance at 60% lower cost via post-training alone

“Terminal-Bench 82.7, up +25.8 from the April preview's 56.9”
moonshotai/Kimi-K3

Simon Willison · Jul 27, 2026

Moonshot AI releases Kimi K3 open weights with restrictive MaaS licensing requiring separate commercial agreements

“If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.”
Inside the Model Factory — Eiso Kant, Poolside AI

Latent Space Blog · Jul 23, 2026

Poolside AI's Model Factory ships models from pre-training to release in eight weeks.

“would rather live in a world with 100 foundation model companies than five even if Poolside were one of the five”
Ornith-1.0: Self-Scaffolding LLMs for Agentic Coding

Simon Willison · Jun 29, 2026

DeepReinforce released Ornith-1.0, an MIT-licensed open-weights agentic coding model built on Gemma 4 and Qwen 3.5.

“Initial impressions are very good - it seems to be able to run the agent harness over many tool calls in a proficient way.”
What's new in the Gemma open model family

Google Developers (Google I/O) · May 22, 2026

Google launches Gemma 4 open model family with four sizes from 2B to 31B parameters

“Our Gemma model, here evaluated on L M Arena, are scoring as well as model 20X the size.”
This Free AI Just Caught The Billion Dollar Giants

Two Minute Papers · Aug 28, 2026

Qwen 3.8 Flash Next, a free open-weights MoE model, rivals paid closed AI systems.

“it seems that we can switch out paid closed systems to open weights AI that we can download and run ourselves forever”
Qwen3.8-Flash-Next

Simon Willison · Aug 26, 2026

Qwen releases 125B MoE model with only 6B active parameters as Qwen4 architecture preview

“a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4”
PipeNetwork/minimax-h3-mlx

Simon Willison · Aug 04, 2026

MiniMax-H3 omni-modal model now runs locally on Apple Silicon via MLX port

“a rainbow colored skunk leaps over a mossy log in a supermarket”
Introducing Gemma 4 models on Amazon Bedrock

AWS Machine Learning Blog · Jun 15, 2026

Google DeepMind's open-weight Gemma 4 model family is now available on Amazon Bedrock.

“Artificial Analysis reports an Intelligence Index of 39 for Gemma 4 31B, well above the median of 15 in the 4B–40B open-weights class.”
AIventure: Vibe Coding Journey

Google Developers (Google I/O) · Jun 12, 2026

Google released AIventure, an open-source game teaching vibe coding and agentic workflows with the Gemma 4 open-weights model.

“you have complete flexibility in how the model is served”