Skip to content
worth noting Open-source

Simon Willison released the open-source evaluation tool smevals for testing LLM models

only one source so far

Simon Willison released the open-source tool smevals for running and grading small evaluation suites across models. An evaluation takes the form of YAML files in a directory, and results can be viewed through a web interface or exported to static HTML.

Simon Willison, an independent commentator and developer following developments around LLMs, released the open-source tool smevals for running small evaluation suites across different model configurations and assessing their results. According to the author, this is the third iteration of his long-running search for a suitable approach to evaluations, which he has been working on for several years, and he says this version finally feels “right" to him.

An evaluation in smevals takes the form of a directory containing YAML files. The tool is launched via uvx smevals, with running an evaluation suite against selected models (using the -m parameter for each model) and grading the results against defined checks handled as separate steps – first run, then grade. Results can be viewed through a local web server using the serve command, or exported to a static HTML page using the build command, which can be hosted anywhere. The tool documentation can be displayed with the command uvx smevals docs.

As an example, the author presented an evaluation suite he built to test the ability of models to write haiku. He said he plans to expand the tool further and use it for his own projects as well.

What changed

Why it matters

The tool offers a simple, scriptable way to systematically compare the behavior of multiple LLM models or prompt versions on the same set of tasks without having to build your own evaluation infrastructure. Separating execution from grading and supporting export to static HTML makes it easy to share and archive results outside the development environment as well.

Two audiences, two different impacts

What this means

01

For individuals

Developers working with LLMs get a freely available tool they can use to test and compare prompt behavior across models without building their own tooling.

What to do Try running your own evaluation suite with the command uvx smevals run and compare the outputs of multiple models.
More practical updates →
02

For a business

Companies developing products built on LLMs can use the tool as a lightweight alternative to building their own evaluation infrastructure for comparing models and configurations.

Development
What to decide Consider trying smevals for systematic A/B testing of prompts and models before deploying to production.
More business impacts →
eval framework LLM testing model evaluation open-source smevals testing tools

Check the original

Event sources

only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.

1
Simon Willison — AI tag (leading independent LLM commentator) community signal · first detected smevals - a small eval suite for evaluating models, prompts, and harnesses