Vals, backed by Andreessen Horowitz, is looking to become the gold standard for AI benchmarking
Vals AI is hoping to make AI benchmarking a more neutral and trustworthy resource in a world increasingly inundated by AI models.

TL;DR
- AI companies use benchmarking for PR and validation, but current systems are often outdated and outsmarted.
- Vals, founded in 2024, is developing a new system for AI benchmarking, focusing on real-world impact and industry-specific tasks.
- The company differentiates itself by not publicly disclosing its test materials, preventing companies from training against them.
- Vals evaluates AI models on their ability to perform complex tasks, produce quality output comparable to humans, and analyzes potential negative implications.
- The startup is expanding its benchmarking capabilities into areas like recursive self-improvement, mental health, cybersecurity, biosecurity, and law of armed conflict.
- Companies pay Vals to test their models, providing crucial data for troubleshooting, improvement, and acquisition decisions.
- Vals has experienced significant growth in revenue and staff since its inception and is expanding its services to federal agencies.
- The company believes its benchmarking system will be central to how AI companies grow, establish public trust, and interact with investors as AI becomes more integrated into the economy.