
What happened
Vals, a 2024-founded AI benchmarking startup, raised $40 million in a series A led by Andreessen Horowitz, after a seed round led by 8VC and Bloomberg Beta.
Why it matters
Vals keeps its test materials undisclosed and scores models on real tasks in law, finance and coding, a pitch aimed at companies that say older benchmarks are being gamed.
What to watch
Its revenue is now eight times last year's level, but that growth hinges on whether paying customers keep buying tests and on its push into federal agency evaluations.
WHO IT HITSIt lands on teams that buy or deploy AI models — procurement and evaluation staff at enterprises and federal agencies — who currently lean on public benchmarks to compare vendors. Vals is pitching those buyers a paid, undisclosed alternative.
Summaries like this, in your inbox every morning.
Vals was formed in 2024 and built its reputation on a simple complaint: benchmarks that are publicly available can be trained against, so they stop proving much. Co-founder Rayan Krishnan says the academic benchmarks were not keeping up with the frontier of new, capable models, and that benchmarks should verify what companies actually advertise. Vals' answer is to keep its test materials undisclosed and to judge models on work tied to specific industries such as law, finance and coding.
The paying-customer model can look counterintuitive, since a company is buying a test it might fail. Krishnan compares it to a student paying the College Board to sit the SAT, and the article frames the value as troubleshooting and improvement over time, with these evaluations increasingly used by companies deciding which AI models to acquire. To date the company has raised a seed round led by 8VC and Bloomberg Beta, then $40 million in a series A led by Andreessen Horowitz; it also recently launched a program providing model evaluations to federal agencies.
The internal signals are mixed in a way worth watching. Vals says revenue is now eight times last year's level, and headcount has tripled from eight to 25 with plans to add 10 to 15 more. Krishnan argues that as AI companies go public and models become a core part of the economy, evaluations of this kind will shape usage and public filings. Whether that holds may depend on how much weight corporate and government buyers place on an undisclosed test whose methods they cannot inspect themselves.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
NYU professor Tristan Buckmaster accused OpenAI of learning his Navier-Stokes work was near a solution, then d…

Compal Electronics is buying out the shares of its listed networking arm, Compal Broadband Networks (CBN), in…

Citi says opposition to data centers has not materially weakened the construction pipeline, even as at least 1…

Pat Gelsinger and Naveen Rao argue in Fortune that AI is creating new jobs as fast as it eliminates old ones…

The Gates Foundation announced a $1 billion commitment to equitable AI tools, with 40% — about $400 million —…

Lucy Guo, who became a billionaire from Scale AI's sale to Meta, said AI is making people work harder, and she…
