smevals - a small eval suite for evaluating models, prompts, and harnesses
dev_tools
Simon Willison is announcing smevals, an evaluation framework for testing and comparing AI models. Built with Jesse Vincent's Prime Radiant research lab, the tool lets developers run evaluation suites across different model configurations, grade results, and explore findings via a web interface or static HTML reports. Willison says this is his third take on evaluation design, and it's now available as open-source software for the development community to use and extend.
Source: https://simonwillison.net/2026/Jul/31/smevals/#atom-everything
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton