The Hallway Track

coding

14 tracked signals on coding.

Introducing GPT-6.1 Sol

OpenAI · OpenAI Blog · Sep 29, 2026

OpenAI launches GPT-6.1 Sol at one-fifth the price of its flagship Astra model

“near-Astra intelligence for coding, computer use, and professional work at one-fifth of Astra's standard API input and output token prices”
GLM-5.3: How Chinese labs keep stride with the frontier

Nathan Lambert · Interconnects · Aug 14, 2026

Z.ai's GLM-5.3 surpasses Claude Fable 5 and GPT-5.6-Sol on select benchmarks

“On many benchmarks the model has surpassed Moonshot AI's Kimi K3 and on some it's surpassed Claude Fable 5 or GPT-5.6-Sol.”
Open-weight AI just hit 2.8 trillion parameters…

Fireship · Jul 22, 2026

Moonshot AI's Kimi K3 is a 2.8T-parameter open-weight model matching frontier closed models on coding benchmarks.

“it has OpenAI and Anthropic terrified because its Trust Me Bro benchmark performance is on par with and in some cases beating Claude Fable and GPT 5.6 Soul”
The New Physics of Business — Garry Tan, Y Combinator

AI Engineer · Jul 17, 2026

YC's Garry Tan claims 400x personal coding productivity gain with AI agents.

“One person does what used to take a thousand people. And I don't mean that as a metaphor. I mean that mechanically this year, the people in this room will do this.”
Grok 4.7 is now available on Amazon Bedrock

AWS Machine Learning Blog · Sep 28, 2026

xAI's Grok 4.7 lands on Amazon Bedrock with 500K context and self-verification for agents

“A model that checks its own output before continuing tends to fail less catastrophically on long trajectories, where an early mistake otherwise compounds through every later step.”
xAI’s Grok 4.6 is now available in Amazon Bedrock

AWS Machine Learning Blog · Sep 21, 2026

xAI's Grok 4.6, a frontier agentic model, is now available in Amazon Bedrock.

“Grok 4.6 widens that surface area considerably: it is available on both the bedrock-mantle and bedrock-runtime endpoints, and it supports the Converse API alongside Chat Completions and Responses.”
Introducing Gemini 3.7 Flash

Google Developers (Google I/O) · Aug 13, 2026

Google launches Gemini 3.7 Flash, its most capable coding and agent workhorse model

“a model that just feels better to build with, landing where you want to go in fewer shots, less back and forth, and with higher fidelity”
Is Coding Solved?

No Priors · Aug 07, 2026

AI lab insiders believe coding will be a solved problem within 6 months

“We're 6 months or towards the end of the year to be completely done with code. Like it's a solved problem.”
Gemini co-leads on project origins and what's next

Google Developers (Google I/O) · May 29, 2026

Google launched Gemini 3.5 Flash, focused on coding and agentic capabilities.

“coding capabilities and agentic experiences are defining what it means to experience AI”
What is Gemini 3.7 Flash?

Google Developers (Google I/O) · Aug 13, 2026

Google launches Gemini 3.7 Flash as its top coding and agent workhorse model

“our most intelligent workhorse model yet for coding and agents”