Everything Is a Rollout — Alex Shaw + Ryan Marten, Terminal-Bench, Harbor, Laude Institute
Harbor framework treats all agent runs as RL rollouts, requiring empirical evaluation like ML models.
“Agentic coding is a form of machine learning. Generated code is best treated as a blackbox artifact whose behavior and generalization should be managed via empirical evaluation like with any ML model.”
Alex Shaw from the Laude Institute introduced Harbor, an agent evaluation and RL environment framework, arguing that agentic coding fundamentally differs from traditional software engineering because code outputs can no longer be predicted before execution. The core thesis—'everything is a rollout'—reframes agent outputs as blackbox ML artifacts requiring empirical measurement rather than static code review. This mental model shift has practical implications for how teams build evaluation infrastructure around AI coding agents.