RL-trained Qwen 27B model Faraday outperforms Claude and GPT-5 on scientific replication tasks
“they can get this so-called AI scientist agent which can outperform you know much larger models such as Claude and GBD5”
9 tracked signals on llm-as-judge.
RL-trained Qwen 27B model Faraday outperforms Claude and GPT-5 on scientific replication tasks
“they can get this so-called AI scientist agent which can outperform you know much larger models such as Claude and GBD5”
Chime built compliance evals for its AI co-pilot Jade by having legal teams co-author risk definitions.
“I would argue that evals are your alignment surface.”
Lyft scaled to seven+ production AI agents at 35% resolution by building an offline-eval quality gate before shipping.
“You don't want to use your users as test data.”
AI engineering evals are fundamentally data science, and teams keep repeating avoidable eval mistakes.
“I'm saying that using LLM judges kind of blindly without validating your validators is bad.”
AWS released Agent-EvalKit, an open-source toolkit for systematically evaluating AI agents' full execution paths.
“An agent might deliver a well-structured, actionable response while hallucinating, fabricating facts because its tools returned empty results.”
Apple's new Evaluations framework lets developers hill-climb and align model judges to reduce drift in AI features.
“This discrepancy between model and human is known as drift, and it is a problem faced by all developers trying to evaluate intelligent features.”
Arize AI presents a 101 framework for evaluating AI agents from tracing to LLM-as-judge meta-evaluation
AWS released the open source Nova Sonic Test Harness to automatically evaluate voice agents at scale without a microphone.
“It runs complete multi-turn conversations with Amazon Nova Sonic automatically, evaluates them using LLM-as-judge techniques, and can even detect cases where the model's audio output doesn't match its text output (audio hallucinations).”
LangChain demos 'Jev,' a fast, cheap System 1 classification model for agent routing, guardrails, and evals.
“System 1 models are a class of AI models built to make fast structured decisions that software can use directly.”