OpenAI launches GPT-6.1 Sol at one-fifth the price of its flagship Astra model
“near-Astra intelligence for coding, computer use, and professional work at one-fifth of Astra's standard API input and output token prices”
14 tracked signals on coding.
OpenAI launches GPT-6.1 Sol at one-fifth the price of its flagship Astra model
“near-Astra intelligence for coding, computer use, and professional work at one-fifth of Astra's standard API input and output token prices”
Z.ai's GLM-5.3 surpasses Claude Fable 5 and GPT-5.6-Sol on select benchmarks
“On many benchmarks the model has surpassed Moonshot AI's Kimi K3 and on some it's surpassed Claude Fable 5 or GPT-5.6-Sol.”
Moonshot AI's Kimi K3 is a 2.8T-parameter open-weight model matching frontier closed models on coding benchmarks.
“it has OpenAI and Anthropic terrified because its Trust Me Bro benchmark performance is on par with and in some cases beating Claude Fable and GPT 5.6 Soul”
OpenAI previews GPT-5.6 Sol, a next-gen model with gains in coding, science, and cybersecurity plus an advanced safety stack.
YC's Garry Tan claims 400x personal coding productivity gain with AI agents.
“One person does what used to take a thousand people. And I don't mean that as a metaphor. I mean that mechanically this year, the people in this room will do this.”
Reflection AI launches Beam, a 501B/23B-active US-trained open MoE model for coding and agentic work
GPT-6.1 Sol is now generally available on Amazon Bedrock with near-Astra intelligence at one-fifth the cost
xAI's Grok 4.7 lands on Amazon Bedrock with 500K context and self-verification for agents
“A model that checks its own output before continuing tends to fail less catastrophically on long trajectories, where an early mistake otherwise compounds through every later step.”
xAI's Grok 4.6, a frontier agentic model, is now available in Amazon Bedrock.
“Grok 4.6 widens that surface area considerably: it is available on both the bedrock-mantle and bedrock-runtime endpoints, and it supports the Converse API alongside Chat Completions and Responses.”
Google launches Gemini 3.7 Flash, its most capable coding and agent workhorse model
“a model that just feels better to build with, landing where you want to go in fewer shots, less back and forth, and with higher fidelity”
AI lab insiders believe coding will be a solved problem within 6 months
“We're 6 months or towards the end of the year to be completely done with code. Like it's a solved problem.”
Google launched Gemini 3.5 Flash, focused on coding and agentic capabilities.
“coding capabilities and agentic experiences are defining what it means to experience AI”
Z.ai's 753B-parameter GLM 5.3 MoE model is now available on Amazon Bedrock for enterprise use
Google launches Gemini 3.7 Flash as its top coding and agent workhorse model
“our most intelligent workhorse model yet for coding and agents”