Agentic Search vs Vector Search for Coding Agents: We Ran the Eval — Braintrust
Systematic evals beat gut-feel releases: 94% pass rates beat 'it looked good'
“I ran 200 different test scenarios, 94% of them passed, so we're releasing this feature.”
A BrainTrust developer engineer presented a framework for why structured evaluations are essential for AI system releases, using OpenAI's April 2025 sycophancy regression as a cautionary example of what happens without them. The talk frames evals as the answer to key questions: model selection, real-world performance across languages and tasks, cost efficiency, and regression detection. The title promises a head-to-head comparison of agentic vs vector search for coding agents, but the provided content is limited to the introductory eval methodology section.