How We Built an Agent That Improves Itself — Zubin Aysola, Weights & Biases
Weights & Biases released Agent Arya, a self-improving research agent using evaluation-driven self-reinforcement.
“tests, assessments, agents, and how you configure them are covariant”
Weights & Biases shipped Agent Arya, a research agent that uses its own Weave monitoring platform to self-reinforce across both offline simulation and production environments. The core technical insight is that agent evaluation is fundamentally covariant — as the agent changes, so must the tests and configurations that measure it, requiring parity between offline and production tracing. This matters because it surfaces a principled evaluation architecture pattern that any team building production agents will face.