A 25-Year-Old Just Raised $40 Million Because AI Benchmarks Are Broken

Vals AI closed a $40 million Series A led by Andreessen Horowitz to build confidential, domain-specific benchmarks for AI models in law, finance, coding, and cybersecurity, betting that secret tests measure real capability better…

Performance analytics charts on a computer screen

Written by Admin Alex · Fact-Checked by M.Ali · Info Verified September 2026

We review and update this article regularly as new information becomes available.

TL;DR: Vals AI, founded by 25-year-old Rayan Krishnan, closed a $40 million Series A led by Andreessen Horowitz to build confidential, domain-specific benchmarks for AI models in law, finance, coding, and cybersecurity. Unlike public leaderboards that models can be tuned to beat, Vals keeps its test material secret and charges companies to evaluate their models against it, similar to how students pay to take the SAT.

Data analytics dashboard on a laptop screen representing AI benchmarking

Ask anyone who works with AI models professionally whether they trust the public benchmarks, and you’ll get a similar answer most of the time: not really. Model makers have gotten good enough at optimizing for known test sets that a leaderboard score increasingly measures how well a lab studied for the test, not how the model performs on work nobody has seen before.

That gap is what Vals AI is betting $40 million on closing. The startup just closed a Series A led by Andreessen Horowitz, on top of an earlier seed round from 8VC and Bloomberg Beta. Founder Rayan Krishnan is 25, a former Palantir intern who also spent time at Microsoft and Stanford’s AI lab before starting the company two years ago.

The model is straightforward once you hear it explained. Vals builds domain-specific tests in law, finance, coding, cybersecurity, and biosecurity, and it keeps the actual test material confidential. Companies pay Vals to run their models against it, the same basic arrangement as a student paying a testing company to sit for the SAT. Krishnan’s pitch is that secrecy is the whole point. A public benchmark becomes a training target the moment it’s published. A private one stays a measurement of actual capability for longer, because nobody can optimize against a test they’ve never seen.

“Evaluation has been done to evaluate intelligence in a very abstract way,” Krishnan said, describing the shift Vals is trying to make. “What we’re doing is actually looking at what are the real impacts.”

The traction numbers back up the pitch, at least so far. Revenue is running eight times what it was a year ago. Headcount has grown from 8 to 25 people, with 10 to 15 more hires planned in the next stretch. That’s fast growth for a company whose entire product is, in a sense, a very expensive and very specific pop quiz.

Why this problem got urgent

Academic benchmarks were never built to track a field moving this fast. A test written to evaluate models in 2024 tells you almost nothing useful about a frontier model released in 2026, and by the time a new academic benchmark gets published, peer-reviewed, and adopted, the models it’s measuring are already a generation or two behind whatever’s actually shipping. Public leaderboards, the kind AI labs love to cite in launch announcements, have the opposite problem. They’re current, but they get gamed constantly, sometimes unintentionally, sometimes not.

Enterprise buyers are the ones actually feeling this gap day to day. A law firm evaluating which model to trust with contract review doesn’t care how a model scores on a general reasoning benchmark. It cares whether the model gets the actual legal work right, consistently, on documents it’s never encountered before. That’s a much narrower and much harder thing to measure, and it’s exactly the kind of test Vals is trying to sell.

Bottom Line

The benchmarking crisis in AI isn’t new, but it’s gotten expensive enough that someone was going to build a business around fixing it. Vals is early, still small at 25 people, and its confidential-test model depends entirely on staying ahead of labs that have every incentive to reverse-engineer what’s being asked. Whether it becomes the SAT of AI evaluation or just one of several competing standards remains to be seen, but the $40 million bet says a16z thinks the market for honest measurement is bigger than the market for another leaderboard.