smevals - a small eval suite for evaluating models, prompts, and harnesses
smevals is a new open-source eval framework for comparing AI model and prompt configurations
“I've been trying to figure out an approach I like for evals for several years now. smevals is my third iteration on the idea and it feels right to me.”
Simon Willison, working with Prime Radiant applied AI research lab, released smevals, a lightweight CLI tool for building and running eval suites against multiple model configurations. The tool separates run, grade, and serve steps, allowing practitioners to test models and prompts with YAML-defined tasks and custom checkers. It addresses a long-standing gap in accessible, structured eval tooling for individual developers and small teams.