Vals.ai
Vals.ai builds independent AI benchmarking infrastructure that evaluates foundation models and agents on private, domain-specific tasks in finance, legal, healthcare, software, and academia, serving AI labs and enterprise teams via a SaaS platform with SDK and enterprise sales.
- Company typePrivate
- Founded2024
- HeadquartersSan Francisco, United States
- Headcount11–50
- GTM typeB2B
- OfferingSoftware
What Vals.ai does
Vals.ai is a San Francisco-based AI evaluation company founded in 2024 that builds independent benchmarking infrastructure for foundation models and agentic AI systems. The company's core product is a SaaS evaluation platform that runs domain-specific benchmarks across finance, legal, healthcare, software engineering, and academic tasks, using private test sets to prevent dataset leakage and expert-authored rubrics to grade performance. The platform supports evaluation of text, code, structured data, and multimodal inputs, and includes agentic capability testing (tool-calling, multi-turn flows, computer-use). Notable benchmarks include the Vals Index (a composite weighted by sectoral contribution to U.S. GDP), Vals Multimodal Index, Legal Research Bench, Harvey's Legal Agent Benchmark, Code Migration, SkillsBench, SWE-bench Verified, Terminal-Bench, and vertical-specific evaluations such as Finance Agent v2, CorpFin v2, MedCode, and MedScribe.
The business model is enterprise SaaS delivered through a self-serve platform with SDK, CLI, and CI/CD integrations for engineering teams, combined with an enterprise sales overlay that provides quote-based pricing, custom evaluations, and private dataset licensing on a case-by-case basis. The company serves AI labs, enterprise engineering teams, and AI application vendors in legal, financial services, healthcare, insurance, and software engineering verticals, with primary buyer pain points centered on selecting and validating AI models for high-stakes production use cases. Vals.ai has earned coverage from Bloomberg, the New York Times, and TechCrunch, achieved SOC 2 Type 2 certification, and partnered with Harvey, Stanford researchers, BenchFlow AI, and industry domain experts to produce benchmarks.
The company was launched with seed funding from Pear VC and operates with a 1-10 employee team. Revenue, headcount growth, and follow-on funding rounds have not been publicly disclosed. The product launch cadence accelerated in 2026 with multiple new benchmark releases, and the company has begun building an open ecosystem through the Valkyrie public benchmark registry.
Vals.ai firmographics
Firmographics- Name
- Vals.ai
- Legal name
- Vals AI, Inc.
- Website
- https://vals.ai
- Company type
- Private
- Founded year
- 2024
- Operating status
- Operating
- Headcount range
- 11–50 employees
- Short description
- Vals.ai builds independent AI benchmarking infrastructure that evaluates foundation models and agents on private, domain-specific tasks in finance, legal, healthcare, software, and academia, serving AI labs and enterprise teams via a SaaS platform with SDK and enterprise sales.
- Ownership category
- akta.pro rank
Vals.ai industry classification
Industry- Product category
- AI Model Evaluation Platform
- NAICS
- Software Publishers (5132), Software Publishers (513210), Other Computer Related Services (541519)
- SIC
- Services-Computer Programming Services (7371), Services-Computer Integrated Systems Design (7373)
- akta.pro primary industry
- AI Observability, Monitoring & Evaluation Platforms (Drift, Quality, Safety) (HDAEANAF)
- akta.pro secondary industries
- Enterprise AI Governance, Risk & Compliance Platforms (Model Risk, Audit, Policies) (HDAEANAE), Audit, Explainability & Accountability Tooling (traceability, reporting) (HDAAAKAL), Responsible AI, AI Governance & Compliance Services (BPAEAHAJ)
Keywords
Where Vals.ai is headquartered
LocationHeadquarters
- HQ city
- San Francisco
- HQ country
- United States
- HQ region
- North America
Markets served
Vals.ai business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Technology or R&D, Personnel, Infrastructure, Marketing or Sales, Operations
Revenue model
- Enterprise Benchmarking Platform: SaaS platform subscription model where enterprises pay for access to evaluation infrastructure, benchmarking tools, and expert review capabilities for assessing AI models and agents in high-value domains.
- Private Benchmark Licensing: Companies license private validation datasets for their own internal validation, with statistical proof correlated to Vals' test suite.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Subscription | Annual | Early Access / Enterprise |
Go-to-market motion2 records
Distribution channels2 records
Marketing channels5 records
Vals.ai product offering
Product offeringCore offering
Vals.ai operates an independent AI model and agent benchmarking platform that evaluates large language models on private, domain-specific tasks spanning finance, legal, healthcare, software engineering, and academia. The platform collects review criteria from subject-matter experts, runs evaluations at scale via SDK/CLI tooling, and publishes indices such as the Vals Index and Vals Multimodal Index. It targets enterprises and AI labs that need unbiased, leakage-resistant performance signals for model selection, regression testing, and deployment decisions.
Product overview
Vals AI operates as a benchmarking and evaluation platform for AI models and agents. The core offering is the Vals AI Platform, an evaluation infrastructure that allows labs and engineering teams to collect review criteria from subject-matter experts, run evaluations at scale, and drive their review process. Built on this platform are industry-specific benchmarks including the Vals Index (aggregated score across domains) and Vals Multimodal Index (finance, coding, education weighted by economic contribution), along with specialized benchmarks for legal research (Legal Research Bench, Harvey's Legal Agent Benchmark), coding (SWE-bench Verified, Vibe Code Bench, Terminal-Bench, ProgramBench, Code Migration), finance (Finance Agent v2, CorpFin v2), and healthcare (MedCode, MedScribe). The company also offers developer tools including an SDK and CLI for CI/CD integration, and maintains the Valkyrie public benchmark registry for community contributions. All benchmarks are designed to be private and secure to prevent dataset leakage, with public validation sets for transparency.
Differentiator
Problem solved
Functional benefit
Products and services
- Vals AI Platform Evaluation infrastructure platform enabling labs and engineering teams to collect review criteria from subject-matter experts, run evaluations of any LLM model at scale, and drive their review process. Supports benchmarking of foundation models and LLM applications on task-specific data.
- Vals Index Aggregated benchmark score combining multiple domain-specific evaluations weighted by economic sector contribution. Top performing models include Claude Fable 5 (75.14%), Claude Opus 4.8 (70.36%), and GPT 5.5 (67.95%).
- Vals Multimodal Index Benchmark measuring AI model performance across finance, coding, and education sectors, weighted by each sector's contribution to the U.S. economy using Federal Reserve Economic Data and Bureau of Labor Statistics.
- Legal Research Bench Private benchmark testing whether agents can handle realistic US legal research across eight practice areas using case-law search, web search, and document retrieval, graded against rubrics authored and peer-reviewed by practicing lawyers.
- Harvey's Legal Agent Benchmark (HLab) Benchmark testing an agent's ability to complete legal work by answering client inquiries using shell and file-editing tools plus skills for Word, Excel, and PowerPoint.
- Code Migration Proprietary benchmark asking whether models can reimplement working programs in another language. Includes CLI split (30 repositories across Python, Java, Kotlin, Rust, and C++) and COBOL-to-Java split.
- SWE-bench Verified Benchmark evaluating software engineering capabilities of AI models on real-world tasks from GitHub repositories.
- Vibe Code Bench Benchmark testing AI model performance on coding tasks, with DeepSeek V4 ranking #1 among open-weight models.
- Terminal-Bench 2.1 Benchmark evaluating terminal and command-line interface capabilities of AI models.
- ProgramBench Benchmark evaluating whether models can reconstruct command-line programs from an executable binary and behavioral specification.
- SkillsBench Benchmark evaluating whether agents improve at software tasks when given reusable, task-specific knowledge. Average score improved from 35.53% without skills to 52.53% with skills.
- Finance Agent v2 Benchmark evaluating AI agent performance on financial tasks and workflows.
- CorpFin v2 Benchmark evaluating AI model performance on corporate finance tasks.
- MedCode Benchmark evaluating AI model performance on medical coding tasks.
- MedScribe Benchmark evaluating AI model performance on medical transcription and documentation tasks.
- Valkyrie Public Benchmark Registry Public benchmark registry and evaluation framework allowing community contribution of benchmarks compatible with Vals AI's evaluation platform. Integrated with SkillsBench.
- Vals AI SDK Python software development kit enabling developers to integrate Vals AI's evaluation infrastructure into their applications for benchmarking LLM performance.
- Vals AI CLI Tools Command-line interface tools for running evaluations, enabling CI/CD integration and automated testing workflows for LLM applications.
Quantifiable outcome
- Claude Fable 5 leads Vals Index at 75.14%, outperforming other publicly available AI technologies by 5% according to Vals AI benchmark tests
- +3 more outcomes
Companies that use Vals.ai
Customer profileSegments6 records
Ideal customer profiles5 records
Vals.ai technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
AI capability8 records
Feature5 records
Vals.ai partnerships and signals
Strategic signalPartnerships
Three partnerships are on record, tiered core and secondary.
- Stanford ResearcherscoreCollaboration with Stanford university researchers to develop domain-specific AI benchmarks and evaluation methodologies.
- Industry ExpertscorePartnership with industry experts to create benchmarks that reflect real-world industry use cases across finance, legal, healthcare, and other domains.
- BenchFlow AIsecondaryPartnered with BenchFlow AI to integrate SkillsBench into Vals AI's Valkyrie evaluation framework, contributing to the public benchmark registry.
Scale indicators3 records
Recent moves6 records
Expansion highlights6 records
Vals.ai competitors and assessment
Company assessmentBroad incumbents
- Hugging Face: Hugging Face hosts the Open LLM Leaderboard and other community benchmarks, directly competing with Vals.ai on independent model evaluation. It also offers a much broader open-source model and dataset ecosystem, making it a broad incumbent in the same evaluation-adjacent space.
- Scale AI: Scale AI operates a large AI evaluation and data-labeling platform serving foundation model labs and enterprises. It overlaps with Vals.ai on model evaluation and benchmarking services, but at much larger scale and with a broader data-services portfolio.
- Weights & Biases: Weights & Biases offers MLOps tooling including LLM evaluation, experiment tracking and reporting. It is a broad incumbent overlapping with Vals.ai on enterprise evaluation, reporting and audit workflows.
- MLPerf (MLCommons): MLCommons runs industry-standard ML benchmarking suites (MLPerf). While focused more on training/inference systems than domain tasks, it is comparable to Vals.ai as a standardized, independent benchmark authority.
- LangSmith (LangChain): LangSmith is an LLM application evaluation, debugging and monitoring suite bundled with the LangChain ecosystem. It competes with Vals.ai on enterprise LLM evaluation and CI/CD-integrated testing workflows.
Direct peers
- Patronus AI: Patronus AI builds enterprise-focused LLM evaluation and safety testing products. It overlaps directly with Vals.ai on domain-specific evaluation (finance, legal) and agent/system-level assessment for enterprise buyers.
- Braintrust: Braintrust provides an evaluation and observability platform for LLM applications with scoring, datasets and CI/CD integration. It is comparable to Vals.ai on the developer-tooling and regression-testing layer for AI applications.
- Artificial Analysis: Artificial Analysis provides independent benchmarks and comparisons of AI models across quality, latency and cost dimensions. It is the closest direct peer to Vals.ai, with similar positioning around unbiased, third-party model evaluation.
- LMSYS (Chatbot Arena): LMSYS Chatbot Arena is a widely cited crowd-sourced benchmark for comparing LLMs. It competes with Vals.ai as an independent evaluation channel cited by major publications and used by labs to validate model launches.
Emerging players
- Stanford CRFM (HELM): Stanford CRFM's HELM benchmark is a leading academic, transparent evaluation of foundation models. It is comparable to Vals.ai as a transparent, multidimensional benchmark, though it operates as a research effort rather than a commercial product.
Market position
Strengths4 records
Weaknesses4 records
Competitive moat4 records
Key risks5 records
Key highlights7 records
Customer concentration
Vals.ai social profiles
Digital presenceVals.ai compliance and trust
Trust signalCompliance1 record
Vals.ai financial estimates
Financial estimateRevenue estimate
Valuation estimate
Vals.ai leadership team
Management profileNumber of profiles
Vals.ai funding detail
Funding detailFunding overview
Funding rounds2 records
Investors6 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Vals.ai M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Vals.ai
What does Vals.ai do?
Vals.ai operates an independent AI model and agent benchmarking platform that evaluates large language models on private, domain-specific tasks spanning finance, legal, healthcare, software engineering, and academia. The platform collects review criteria from subject-matter experts, runs evaluations at scale via SDK/CLI tooling, and publishes indices such as the Vals Index and Vals Multimodal Index. It targets enterprises and AI labs that need unbiased, leakage-resistant performance signals for model selection, regression testing, and deployment decisions.
Is Vals.ai a public or private company?
Vals.ai is a private company. It is classified as venture growth investor backed and is currently operating.
When was Vals.ai founded?
Vals.ai was founded in 2024. It employs 11 to 50 people.
Where is Vals.ai based?
Vals.ai is headquartered in San Francisco, United States, in the North America region.
How does Vals.ai make money?
Two revenue lines are on record. Enterprise Benchmarking Platform is the primary driver. The others are private Benchmark Licensing.
Who are Vals.ai's main competitors?
Broad incumbents on record are Hugging Face, Scale AI, Weights & Biases, MLPerf (MLCommons) and LangSmith (LangChain). Direct peers are Patronus AI, Braintrust, Artificial Analysis and LMSYS (Chatbot Arena). Stanford CRFM (HELM) is listed as an emerging player.
Does Vals.ai have an API?
Yes. Vals AI provides an SDK and API for evaluating LLM models and applications. The platform allows labs and engineering teams to collect data, run evaluations at scale, and drive their review process. Documentation is available at docs.vals.ai. Developer documentation is at docs.vals.ai/get_started/introduction.
What industry is Vals.ai in?
Vals.ai's product category is AI Model Evaluation Platform. Its primary akta.pro industry code is HDAEANAF, AI Observability, Monitoring & Evaluation Platforms (Drift, Quality, Safety), with a secondary code of HDAEANAE, Enterprise AI Governance, Risk & Compliance Platforms (Model Risk, Audit, Policies). Its NAICS code is 5132 and its SIC code is 7371.