Simon Willison released the open-source evaluation tool smevals for testing LLM models
Simon Willison released the open-source tool smevals for running and grading small evaluation suites across models. An evaluation takes the form of YAML files in a directory, and results can be viewed through a web interface or exported to static HTML.
Simon Willison, an independent commentator and developer following developments around LLMs, released the open-source tool smevals for running small evaluation suites across different model configurations and assessing their results. According to the author, this is the third iteration of his long-running search for a suitable approach to evaluations, which he has been working on for several years, and he says this version finally feels “right" to him.
An evaluation in smevals takes the form of a directory containing YAML files. The tool is launched via uvx smevals, with running an evaluation suite against selected models (using the -m parameter for each model) and grading the results against defined checks handled as separate steps – first run, then grade. Results can be viewed through a local web server using the serve command, or exported to a static HTML page using the build command, which can be hosted anywhere. The tool documentation can be displayed with the command uvx smevals docs.
As an example, the author presented an evaluation suite he built to test the ability of models to write haiku. He said he plans to expand the tool further and use it for his own projects as well.
Why it matters
The tool offers a simple, scriptable way to systematically compare the behavior of multiple LLM models or prompt versions on the same set of tasks without having to build your own evaluation infrastructure. Separating execution from grading and supporting export to static HTML makes it easy to share and archive results outside the development environment as well.
Two audiences, two different impacts
What this means
For individuals
Developers working with LLMs get a freely available tool they can use to test and compare prompt behavior across models without building their own tooling.
For a business
Companies developing products built on LLMs can use the tool as a lightweight alternative to building their own evaluation infrastructure for comparing models and configurations.
DevelopmentCheck the original
Event sources
only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.