deniz.in

Markets

Weather

Loading weather

· via TechCrunch

a16z-backed Vals raises $40M to fix AI benchmarking with private, task-based tests

Startup Vals, which just raised a $40M Series A led by Andreessen Horowitz, keeps its test materials secret and grades models on real industry work, betting evaluations will underpin public trust in AI.

a16z-backed Vals raises $40M to fix AI benchmarking with private, task-based tests

Vals, a San Francisco startup founded in 2024, has raised $40 million in a Series A round led by Andreessen Horowitz, according to TechCrunch, with the stated ambition of becoming the AI industry's standard-setter for evaluation. The round, closed last month, follows a seed round last year led by 8VC and Bloomberg Beta, and it arrives alongside rapid growth: the company says its revenue is now eight times what it was a year ago.

The pitch targets a weakness the AI industry mostly acknowledges. Benchmarks have become the default way labs validate their models and, whenever the numbers swing their way, advertise superiority over competitors. But many of those tests predate today's models, and because most benchmark sets are public, labs can train on them — which makes a strong score about as meaningful as acing an exam after reading the answer key.

Keeping the test papers secret

Vals' central structural difference is that it does not disclose its test materials, closing off the easiest path to contamination. Instead of measuring abstract knowledge — whether a model knows enough to pass a bar-exam-style test — the company evaluates whether models can complete complex, domain-specific tasks in fields such as law, finance and coding.

Co-founder Rayan Krishnan, who is 25, previously interned at Palantir and, as a Stanford undergraduate, worked at Microsoft and the university's AI lab. He told TechCrunch that Vals grew out of watching new, highly capable models arrive while "the academic benchmarks [were] not keeping up with that frontier advance." The question he wants answered instead is whether models "can do work that produces a product of the same quality as a human within every domain."

Vals also tests for failures, not just capabilities. Krishnan said the company analyzes what "the negative implications would be" if models "ran wild in the world." Its coverage has expanded beyond traditional industries into less conventional territory, including recursive self-improvement, mental health, cybersecurity, biosecurity, and whether models can correctly apply the law of armed conflict under the Geneva Convention.

Paying for the grade

The business model sounds counterintuitive at first: AI companies pay Vals to test them, even though the result could be an unflattering score. Krishnan compares the arrangement to a student paying the College Board to sit the SAT — the measurement itself is the product, because it shows companies where their models fall short and whether they improve over time. According to TechCrunch, these evaluations are already becoming a key input into decisions by businesses choosing which AI models to adopt.

The headcount reflects the growth. Vals started the year with eight employees, has tripled to 25, and plans to hire another 10 to 15 people while relocating to a larger office, Krishnan told TechCrunch. The company has also launched a program that provides model evaluations to federal agencies.

Why it matters

Benchmarks shape purchasing decisions and public perception of AI, and the current system is demonstrably gameable. A private, undisclosed test suite fixes the contamination problem but concentrates trust in a single company whose materials cannot be independently inspected — whoever becomes the "gold standard," as TechCrunch puts it, effectively becomes the industry's referee. Krishnan argues that formal evaluations are where the industry is heading regardless. With AI firms moving toward public markets — he cites SpaceX's listing, an Anthropic IPO slated for later this year and a likely OpenAI debut — he expects the kind of benchmarks Vals builds to drive model usage and become "a central part of how these companies submit public filings" and frame their AI investments to shareholders.

  • #ai-benchmarks
  • #evaluation
  • #startups
  • #venture-capital
  • #ai-safety

Related posts