
Vals AI raised $40M from a16z to grade LLMs on real business tasks like financial analysis and coding, not just academic tests. Its Vals Index ranks models by actual work performance.
Alpha Score of 74 reflects strong overall profile with strong momentum, moderate value, strong quality, strong sentiment.
Vals AI has raised $40 million in a Series A round led by Andreessen Horowitz to build independent evaluation benchmarks for AI models, targeting the growing gap between how language models perform on standardized tests and how they actually behave when doing economically valuable work.
The funding marks a significant leap for a company that was bootstrapped with an estimated annual recurring revenue of $1.3 million as recently as late 2025.
Instead of testing whether an AI can solve abstract logic puzzles or complete sentences from Wikipedia, the company evaluates frontier LLMs on tasks that businesses actually pay for: financial analysis, coding, legal research, and web search.
The company's flagship product, the Vals Index, aggregates performance data across these real-world categories and ranks models accordingly. The most recent update to the index, from August 2026, showed Claude Fable 5 sitting at the top with a score of 75.14%.
Vals AI has developed several specialized benchmarks under its umbrella, including the Finance Agent Benchmark and the Web Search Index. Both are built in collaboration with domain experts, not just ML researchers.
Most widely cited benchmarks were designed for academic contexts and have increasingly become targets that model developers optimize for directly. This creates a real headache for enterprise buyers trying to choose between models. When OpenAI, Anthropic, Google, and Meta all claim state-of-the-art performance on overlapping benchmarks, the numbers stop being useful for procurement decisions.
Andreessen Horowitz has been one of the most aggressive venture investors in AI infrastructure, backing companies across the stack from chip design to application layers. Adding Vals AI to the portfolio signals a belief that the evaluation layer of the AI ecosystem is underfunded relative to its importance.
The $40 million should give Vals AI room to expand its benchmark coverage into new verticals, hire domain experts, and build out the tooling that enterprise customers need to run evaluations on their own proprietary data. The company has also announced new products alongside the funding, suggesting a push to move beyond benchmarking as a standalone offering and into continuous model monitoring.
Drafted by a large language model from the source reporting linked above, then screened by automated publishing checks. It is not read by a journalist before publication. Some articles cite our Alpha Score. Verify prices and figures against the original source. Educational coverage, not personalized advice.