Skip to content
context Business and investment

Startup Vals raises 40 million dollars from Andreessen Horowitz for industry-specific AI benchmarks

only one source so far

Startup Vals raised 40 million dollars in a Series A round led by Andreessen Horowitz to develop AI benchmarks that test the ability of models to handle complex industry-specific tasks (law, finance, code) and, unlike traditional tests, do not make test materials publicly available.

Startup Vals, founded in 2024, raised 40 million dollars in a Series A round led by Andreessen Horowitz, according to information provided by the company. The previous seed round was led by 8VC and Bloomberg Beta. The company positions itself as an alternative to established benchmarking systems, which, according to co-founder Rayan Krishnan, cannot keep pace with the development of modern models and measure general knowledge (e.g. through tests similar to entrance exams) more than practical utility.

Unlike traditional benchmarking, Vals does not make specific test materials publicly available, to prevent models from being trained directly on the tests. Instead of general knowledge, the company evaluates whether models can handle complex tasks in specific fields — including law, finance and programming — at a level comparable to human work, including assessments of potential negative impacts. According to Krishnan, the company is also expanding testing into areas such as recursive self-improvement of models, mental health, cybersecurity, biosafety and application of the Geneva Convention.

The business model involves companies paying Vals to test their own models, much as a student pays to take the SAT — according to Krishnan, the results then inform decisions about deploying or acquiring models. The company says that revenue is currently eight times last year's level and that its employee count has grown from eight to 25, with plans to hire another 10 to 15 people and move to larger headquarters. Vals also recently launched a program offering model evaluations to federal agencies.

Founder Rayan Krishnan (25 years old) previously interned at Palantir and worked for Microsoft and the university's artificial intelligence laboratory while studying at Stanford. According to him, similar benchmarks will also play a role in how AI companies going public (he mentions SpaceX, the planned IPO of Anthropic and the expected IPO of OpenAI) communicate their investments and model capabilities in public filings.

What changed

Why it matters

Companies purchasing or deploying AI models in specialized fields gain an additional independent source of evaluation that, according to its creator, better reflects the real-world work performance of models than older academic tests, and could therefore influence decisions about choosing a supplier as well as how companies communicate with investors.

Two audiences, two different impacts

What this means

01

For individuals

Professionals in fields such as law, finance or programming can expect their work to serve as a benchmark for comparing the quality of AI model outputs with human work.

What to do When assessing the capabilities of AI tools in a specific field (law, finance, coding), give more weight to industry-specific benchmarks than to general tests.
More practical updates →
02

For a business

Companies purchasing or deploying AI models gain another independent way to verify whether a model can handle industry-specific tasks (law, finance, coding) at a level comparable to human work, which, according to Vals, influences decisions about acquiring models; the company now also offers the service to US federal agencies.

Risks and compliance
What to decide Consider using independent industry-specific benchmarks (e.g. Vals) as an additional source of information when deciding whether to deploy or purchase AI models for specific fields.
More business impacts →
AI benchmarking AI evaluace Andreessen Horowitz financování Series A Vals

Check the original

Event sources

only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.

1
TechCrunch AI independent context · first detected Vals, backed by Andreessen Horowitz, is looking to become the gold standard for AI benchmarking