Google DeepMind launches Gemini 4 Argon with industry-first 1M output tokens and SOTA benchmarks
benchmarks
44 tracked signals on benchmarks.
Anthropic's Claude Opus 5 matches Fable performance at half the price
“Anthropic's Claude Opus 5 launch triggered a mix of benchmark scrutiny, strong anecdotal coding-agent praise, and renewed debate about frontier model evaluation.”
Anthropic made Claude Fable 5, a Mythos-class model at least 2x Opus's size, generally available to everyone.
“It is a feat of incredible engineering (and commitment to access) to make these research models GA, and the benchmarks are great… with asterisks.”
Anthropic models now hold the top three Agent Arena spots as Sonnet 5.5 debuts at #3
“good, cheap AND fast”
Claude Opus 5.5 ships, leads SimpleBench at 88.4% and dominates explainer video creation
Qwen 3.8 27B matches GPT-5.6 Luna and nearly ties 1.7T-parameter DeepSeek on benchmarks
“Qwen 3.8 27B is a truly astonishing model.”
AI now completes programming tasks taking humans weeks, in hours via MirrorCode benchmark
“We also found that AI models are improving rapidly over time. Leading models from a year ago would have scored about 30%, and were limited to simpler programs, such as a calendar utility.”
Traditional single-number benchmarks fail to capture that modern AI capability scales with test-time compute budget.
“The capability of the model is a function of how much money you put into it.”
GLM-5.2 is now the leading open weights LLM on Artificial Analysis's Intelligence Index, released under MIT license.
“GLM-5.2 is the leading open weights model on the Intelligence Index v4.1.”
Microsoft released MAI-Thinking-1, a reasoning model trained without third-party distillation, with a 109-page transparency-heavy report.
“hillclimbed from scratch”
Frontier AI models score below 50% on new agentic enterprise IT benchmark ITBench-AA
AI has progressed from grade-school math to contributing to unsolved Navier-Stokes research
“AI systems have gone from struggling with grade-school math to helping solve research problems that have resisted mathematicians for decades, including Navier-Stokes.”
NVIDIA's SWE-Serve benchmark tests whether AI coding agents' patches work in live model serving, not just local tests.
“An AI coding agent’s patch can pass tests yet fail when the server loads a real model and handles requests.”
Manually reviewing 100-1,000 real examples beats any automated benchmark for evaluating AI models.
“there's just no better eval than looking at 100 examples or 1,000 examples”
Together AI built a benchmark to test whether frontier LLMs can generate efficient multi-GPU kernels
“we've sort of shifted the bottleneck to multi-GPU communication”
Harvey built a domain-specific research lab by leveraging frontier ecosystem rather than competing with it directly.
“it's an unfair game competing with the frontier labs if you're an application layer company.”
Qwen 3.8 Max challenges OpenAI and Anthropic at 5-10x lower API pricing
“it sat there thinking for 16 days, starting from an empty folder, writing, testing, and repairing its own code”
AI labs optimize for benchmark scores rather than real-world model quality, warns industry insider
“unfortunately the teams are not getting better models overall but better Elm Marina models whatever that is possibly something with a lot of nested list bullet points and emojis”
DeepSeek V4-Flash 0731 matches GPT-5.6 performance at 60% lower cost via post-training alone
“Terminal-Bench 82.7, up +25.8 from the April preview's 56.9”
Current AI safety policies fail to account for test-time compute, where model capability scales with money spent.
“The capability of the model is a function of how much money you put into it, basically.”
Role-playing persona AI systems are being deployed as civic and pedagogical infrastructure, but persona evaluations may be fundamentally flawed.
“Mine is not better by default. The one thing it is is open. You can read every line of what shapes the persona.”
GLM-5.2 passes the 'frontier model that happens to be open' vibe check with credible out-of-sample validation.
“this is a frontier model that just happens to be open”
NVIDIA Blackwell delivered a clean sweep in MLPerf Training v6.0, leading on scale and per-accelerator performance.
“NVIDIA delivered a clean sweep in MLPerf Training v6.0, the latest edition of industry-standard AI training benchmarks developed by the MLCommons consortium.”
Cognition launches FrontierCode, a benchmark grading code quality and maintainability over passing-test 'slop'.
“Many SWE-bench-Passing PRs Would Not Be Merged into Main”
A new benchmark, SocioHack, shows RL-trained LLMs can discover loopholes that game society's rule systems while staying formally compliant.
“an RL-trained model discovers strategies that remain formally compliant, yet undermine the intended purpose of those systems”
Databricks hosted the inaugural Grounded Reasoning Cup to evaluate AI agents live
Real-world inference data—not benchmarks—is the key signal for continual learning at scale.
“the real unlock moving forward will be continual learning”
Current LLM benchmarks fail to measure learning ability across sequential tasks over time.
“imagine that every time you do something, you completely forget your memory... That's the premise under which we're evaluating language models today.”
LatchBio operates as a vertical AI lab building benchmarks and agents for biology's data explosion
“We are basically a vertical AI lab for benchmark and agent engineering.”
DeepSWE is a contamination-resistant coding benchmark with 113 original tasks across ~100 repositories
Snorkel AI argues every company needs a custom benchmark built from production agent traces.
“It's not a static benchmark. It's a constantly populated data set from your production traces.”
New cybersecurity benchmark exposes AI world-modeling gap with 1-2% model success rates
“models have one to two% success rate on this generic benchmark and the reason is the current model even though they're really good they can't really build a dynamic model of what's happening in the world”
NVIDIA leads agentic coding performance on AA-AgentPerf, the industry's first multi-vendor agentic AI benchmark.
“Artificial Analysis AgentPerf (AA-AgentPerf) offers the industry's first multi-vendor open benchmarks profiling trajectories that are representative of real-world AI agent coding tasks.”
Evals are broken and often misleading, but engineers should still use them well in agentic workflows.
“evals are broken and you should use them anyway”
Hugging Face expands EVA-Bench to 3 domains, 121 tools, and 213 scenarios for agent evaluation.
NVIDIA Blackwell sets a new STAC-AI benchmark record for LLM inference in financial services
Mistral Large 4 released, benchmarks criticized as saturated by frontier models
“The benchmark is saturated. Frontier models are tested with an armadillo in fishnet tights jaywalking on Mars.”
Long-horizon for AI agents is a scalar metric, not a binary category, and shifts rapidly.
“long horizon is really kind of a scalar metric”
AI benchmark prompts are so unrealistic that experienced engineers would never write them in practice
“No one writes prompts like these ever.”
Databricks claims lakehouse architecture outperforms traditional data warehouses in 2026 benchmarks
Hugging Face published research on benchmark optimization in speech recognition
Systematic study finds no evidence AI labs train models to draw pelicans on bicycles better
“Pelicans aren't drawn any better than other animals. Bicycles aren't drawn any better than other vehicles. And no lab draws the combination better than its pelicans and bicycles already predict.”
Hugging Face introduces TutorMoments benchmark for AI tutoring intervention timing
No content was provided to extract a signal from.