Vals, backed by Andreessen Horowitz, is looking to become the gold standard for AI benchmarking
First reported by TechCrunch ·
You can now pay for independent AI model evaluations that verify real-world capabilities beyond theoretical knowledge.
Vals, an AI benchmarking startup founded in 2024 and backed by Andreessen Horowitz, has raised $40 million in a Series A funding round. The company aims to create a new standard for evaluating AI models, addressing limitations in existing academic benchmarks that struggle to keep pace with rapid industry advancements. Vals differentiates itself by not publicly disclosing its test materials, preventing companies from training models to "cheat" on evaluations. Instead, it focuses on assessing AI models' ability to perform complex, industry-specific tasks, such as those in law and finance, and also considers potential negative implications. The startup's revenue has grown eightfold in the past year, and its team has expanded from eight to 25 employees. Vals is also developing benchmarks for areas like recursive self-improvement, mental health, cybersecurity, biosecurity, and international law, and has launched a program to provide evaluations to federal agencies.
The significant funding and focus on private, task-specific benchmarks signal a market shift towards more rigorous and practical AI evaluation, moving beyond academic exercises. This approach is crucial as AI models become integrated into critical sectors like law and finance, where their real-world performance and potential negative impacts require robust scrutiny. Vals's strategy suggests a growing demand for objective third-party validation to build trust and facilitate responsible AI deployment.
Companies that develop or deploy AI models will increasingly face pressure to undergo these sophisticated, proprietary evaluations. This could become a key differentiator in a competitive market, influencing investment decisions and regulatory compliance. The focus on potential negative outcomes also suggests a future where AI safety and risk assessment are as critical as performance metrics, shaping product development and public perception.
AI-written summary. May contain errors.