BenchFlow
BenchFlow is an open-source research lab building AI agent evaluation infrastructure — SkillsBench benchmark, ClawsBench simulated workplaces, and a hardened agent simulation runtime — used internally at major AI labs including Google DeepMind, Meta, and Microsoft.
- Company typePrivate
- Founded2024
- HeadquartersSan Francisco, United States
- Headcount1–10
- GTM typeB2B
- OfferingSoftware
What BenchFlow does
BenchFlow (legal entity Benchmarkthing Inc.) is an open-source frontier environment lab for AI agents, founded in late 2024 and headquartered in San Francisco at Founders, Inc. (2 Marina Blvd). The company builds the environments and benchmarks used to evaluate and train AI agents on realistic computer work rather than static prompts. Its portfolio comprises three interconnected open-source projects: BenchFlow Runtime (an agent simulation runtime with a hardened sandbox against reward-hacking and full trajectory capture), SkillsBench (a benchmark of 86–94+ procedural skill tasks across 11 professional domains), and ClawsBench (five high-fidelity mock workplaces — Gmail, Calendar, Docs, Drive, Slack — wire-compatible with the real Google Workspace and Slack APIs). The runtime executes verified ACP agents including Claude Code, Codex, Gemini CLI, OpenCode, OpenHands, OpenClaw, and Pi across single-agent, multi-agent, and multi-round evaluation patterns, with structured reward output and TOML task definitions.
BenchFlow operates a community-led, open-source business model rather than a traditional SaaS pricing motion. All three core projects are freely available under permissive licenses on GitHub and HuggingFace, and the company has generated $50K+ in revenue from benchmark hosting and related activities prior to "agent-eval" becoming a named category. Its primary users are frontier AI research labs — Google DeepMind (which quality-filtered SkillsBench down to a 60-task internal set), Meta, Microsoft, Tencent, GLM, and Minimax — together with academic researchers at Stanford, CMU, UC Berkeley, Ohio State, UC Santa Cruz, Dartmouth, Boston University, and UNC. The go-to-market runs through Discord (50 to 600 members in eight weeks), two hackathons (600+ participants, $100K+ in prize pools), the Agent Skills '26 workshop at ACM CAIS, and arXiv-published research (SkillsBench 2602.12670; ClawsBench 2604.05172, cited in the Qwen 3.6 Max model card).
The company has raised $1M in seed funding (February 2025) from a16z Scout Fund, Eastlink Capital, and Founders, Inc., employs 1–10 full-time staff, and operates with a 40-plus-person distributed research team. Cumulative platform output includes 100K+ trajectories, 60+ hosted benchmarks (SWE-Bench, WebArena, Terminal-Bench among them), and 7,308 SkillsBench public-sweep trajectories alongside 7,224 ClawsBench evaluation trials.
BenchFlow firmographics
Firmographics- Name
- BenchFlow
- Legal name
- Benchmarkthing Inc.
- Website
- https://benchflow.ai
- Company type
- Private
- Founded year
- 2024
- Operating status
- Operating
- Headcount range
- 1–10 employees
- Short description
- BenchFlow is an open-source research lab building AI agent evaluation infrastructure — SkillsBench benchmark, ClawsBench simulated workplaces, and a hardened agent simulation runtime — used internally at major AI labs including Google DeepMind, Meta, and Microsoft.
- Ownership category
- akta.pro rank
BenchFlow industry classification
Industry- Product category
- AI Agent Evaluation Software
- akta.pro primary industry
- AI Benchmarking, Profiling & Performance Monitoring (HDAAAAAJ)
- akta.pro secondary industries
- Agents & Autonomous Workflows (Tool Use, Planning, Multi-Agent) (HDAAACAF), Model Serving, Inference & Deployment Platforms (APIs, Edge/On-Prem) (HDAEANAC)
Keywords
Where BenchFlow is headquartered
LocationHeadquarters
- HQ city
- San Francisco
- HQ country
- United States
- HQ region
- North America
Offices1 record
Markets served
BenchFlow business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Technology or R&D, Personnel, Infrastructure, Marketing or Sales, Operations
Revenue model
- Research/Open Source: BenchFlow operates as an open-source research organization. All three core projects (SkillsBench, ClawsBench, BenchFlow runtime) are freely available on GitHub and HuggingFace. The company appears to be funded through research grants and academic partnerships rather than product revenue.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Freemium | Others | Free open source tier |
Go-to-market motion1 record
Distribution channels1 record
Marketing channels7 records
BenchFlow product offering
Product offeringCore offering
BenchFlow is a frontier environment lab for AI agents that ships three interconnected open-source projects: SkillsBench (a benchmark of 86–94+ tasks across 11 professional domains evaluating whether procedural skills improve agent performance), ClawsBench (five wire-compatible mock workplaces — Gmail, Calendar, Docs, Drive, Slack — for rigorous agent capability and safety evaluation), and the BenchFlow runtime (a sandboxed agent simulation platform supporting single-agent, multi-agent, and multi-round evals with hardened verifiers and full trajectory capture).
Product overview
BenchFlow is a frontier environment lab for AI agents that provides a unified platform for building and evaluating AI agent environments. The portfolio consists of three connected products: (1) BenchFlow Runtime—the core SDK and cloud runtime that executes agent evaluations in sandboxed environments with hardened verifiers; (2) SkillsBench—a benchmark evaluating whether procedural skills improve agent performance across 86-94+ tasks in 11 domains; and (3) ClawsBench—five high-fidelity mock workplace environments (Gmail, Calendar, Docs, Drive, Slack) wire-compatible with real APIs for safety-evaluable agent testing. The runtime serves as the execution engine powering both benchmarks, enabling single-agent, multi-agent, and multi-round evaluation patterns. All three projects are open source.
Differentiator
Problem solved
Functional benefit
Brands
- SkillsBench: The first benchmark for whether procedural skills make agents better at real work. 86 tasks, 11 domains.
- ClawsBench
- BenchFlow Runtime
Products and services
- SkillsBench The first benchmark for evaluating whether procedural skills — instructions, scripts, and references that agents load on demand — make AI agents better at real work. Contains 86–94+ tasks across 11 professional domains with outcome-based verifiers. Used internally by Google DeepMind, Meta, Microsoft, Tencent, GLM, and Minimax.
- ClawsBench Five high-fidelity mock workplaces (Gmail, Calendar, Docs, Drive, Slack) wire-compatible with upstream Google Workspace and Slack APIs. Enables rigorous evaluation of agent capability and safety in stateful, multi-service workflows without risking real service data. Covers 44 structured tasks including single-service, cross-service, and safety-critical scenarios.
- BenchFlow Runtime The agent simulation runtime that powers SkillsBench and ClawsBench. Provides a unified SDK and cloud platform for running agent evaluations with one Scene-based lifecycle supporting single-agent, multi-agent, and multi-round evals. Sandboxed and hardened against reward hacking with full trajectory capture. Compatible with ACP agents (Claude Code, Codex, Gemini CLI, OpenCode, OpenHands, OpenClaw, Pi).
Quantifiable outcome
- Human-curated skills increased task completion rates by 16.2 percentage points on average in SkillsBench
- +3 more outcomes
Companies that use BenchFlow
Customer profileNamed customers6 records
Segments3 records
Ideal customer profiles3 records
BenchFlow technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- No
- API docs
- API detail
Core technology
AI maturity
App detail
Integration9 records
AI capability4 records
Feature4 records
BenchFlow partnerships and signals
Strategic signalPartnerships
19 partnerships are on record, tiered core and minor.
- ACM CAIScoreAgent Skills '26 workshop accepted as one of five workshops at ACM CAIS 2026. BenchFlow organized this inaugural workshop on agent skills with live SkillsBench design challenges.
- RLWRLDminorCollaborator on ClawsBench development alongside the core BenchFlow team. Contributed to the research and implementation of the high-fidelity mock workplace environments.
- Ohio State UniversityminorAcademic collaborator on ClawsBench research project, contributing to the study of agent capability and safety in simulated workspaces.
- Stanford UniversityminorAcademic collaborator on ClawsBench project, part of multi-institution research effort on AI agent evaluation.
- Carnegie Mellon UniversityminorAcademic collaborator on ClawsBench project, contributing research expertise to agent capability and safety evaluation.
- UC BerkeleyminorAcademic collaborator on ClawsBench project, part of the research team studying AI agent evaluation methodologies.
- AmazonminorIndustry collaborator on ClawsBench research project, part of the 40-person research team including affiliates from Amazon.
- UC Santa CruzminorAcademic collaborator on ClawsBench project.
- Dartmouth CollegeminorAcademic collaborator on ClawsBench project.
- Boston UniversityminorAcademic collaborator on ClawsBench project.
- UNC (University of North Carolina)minorAcademic collaborator on ClawsBench project.
- AGI HousecoreHost venue for Agent Skills Hackathon II, which combined with Hackathon I to draw 600+ participants with over $100K in prize pools.
- Amazon (SkillsBench research)minorIndustry affiliate on the SkillsBench research team (40 researchers) studying whether AI agents can effectively teach themselves new skills.
- Anthropic (SkillsBench research)minorIndustry affiliate on the SkillsBench research team.
- ByteDance (SkillsBench research)minorIndustry affiliate on the SkillsBench research team.
- Foxconn (SkillsBench research)minorIndustry affiliate on the SkillsBench research team.
- Google (SkillsBench research)minorIndustry affiliate on the SkillsBench research team.
- OpenAI (SkillsBench research)minorIndustry affiliate on the SkillsBench research team.
- Founders, Inc.coreHost location for Agent Skills Hackathon I and coworking sessions. Xiangyi works out of Founders, Inc. at 2 Marina Blvd, San Francisco.
Scale indicators10 records
Recent moves6 records
Expansion highlights6 records
BenchFlow competitors and assessment
Company assessmentDirect peers
- Braintrust: Braintrust provides an enterprise LLM evaluation platform with built-in scorers, datasets, and experiment tracking. It targets the same agent-evaluation use cases as BenchFlow's SkillsBench/ClawsBench, focused on production-grade agent quality measurement rather than open research.
- LangChain (LangSmith): LangSmith is LangChain's evaluation and observability platform for LLM agents, offering tracing, dataset management, and evaluators. It directly competes with BenchFlow in the agent-evaluation tooling space, though it is positioned as a broader dev platform rather than a pure benchmark/environment lab.
- Patronus AI: Patronus AI builds evaluation and safety testing products for LLM applications, including agent and RAG evaluation suites. It addresses the same core problem as BenchFlow (rigorous, reproducible agent assessment) but via a paid commercial offering rather than open-source benchmarks and environments.
- Galileo (Galileo AI): Galileo offers an LLM evaluation and observability platform with metrics for hallucination, safety, and agent quality. Like BenchFlow, it serves AI teams that need rigorous agent evaluation, but it is a commercial SaaS product rather than open-source infrastructure.
Emerging players
- Anyscale: Anyscale (built on Ray) provides infrastructure and tooling for scalable AI workloads, including RLHF and agent evaluation pipelines. It serves overlapping AI engineering teams but is focused on distributed compute rather than agent benchmark design and mock environments.
- Confident AI (DeepEval): Confident AI maintains DeepEval, an open-source LLM evaluation framework with metrics for answer relevance, hallucination, and agent quality. It is a smaller-scale open-source peer to BenchFlow, focused on component-level evaluation rather than full workplace simulation.
Broad incumbents
- Weights & Biases: W&B is an established ML experiment tracking and evaluation platform used widely by frontier AI labs for training and eval workflows. It overlaps with BenchFlow on agent experiment tracking but as part of a much broader ML lifecycle offering rather than a focused agent benchmark/environment suite.
- Hugging Face: Hugging Face hosts datasets, models, and benchmark Spaces (including SkillsBench and ClawsBench on benchflow's HF org). It is the primary distribution channel for open-source AI assets and an indirect competitor in providing community-driven evaluation infrastructure, though it does not own its own benchmarks.
- OpenAI (Evals framework): OpenAI released its open-source Evals framework for benchmarking LLMs and maintains substantial internal evaluation infrastructure. As a frontier lab with proprietary benchmarks and one of the labs affiliated with BenchFlow's research, it shapes the ecosystem in which BenchFlow's benchmarks are (or are not) adopted as standards.
- Anthropic (Claude Evals): Anthropic operates its own internal Claude evaluation framework and uses SkillsBench internally. As both a customer and a frontier lab with proprietary evals, Anthropic is a broad incumbent whose internal tooling and product direction directly shape BenchFlow's addressable market.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat6 records
Key risks5 records
Key highlights7 records
Customer concentration
BenchFlow social profiles
Digital presenceBenchFlow financial estimates
Financial estimateRevenue estimate
Valuation estimate
BenchFlow leadership team
Management profileNumber of profiles
Profiles3 records
BenchFlow funding detail
Funding detailFunding overview
Funding rounds2 records
Investors4 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
BenchFlow M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about BenchFlow
What does BenchFlow do?
BenchFlow is a frontier environment lab for AI agents that ships three interconnected open-source projects: SkillsBench (a benchmark of 86–94+ tasks across 11 professional domains evaluating whether procedural skills improve agent performance), ClawsBench (five wire-compatible mock workplaces — Gmail, Calendar, Docs, Drive, Slack — for rigorous agent capability and safety evaluation), and the BenchFlow runtime (a sandboxed agent simulation platform supporting single-agent, multi-agent, and multi-round evals with hardened verifiers and full trajectory capture).
Is BenchFlow a public or private company?
BenchFlow is a private company. It is classified as venture growth investor backed and is currently operating.
When was BenchFlow founded?
BenchFlow was founded in 2024. It employs 1 to 10 people.
Where is BenchFlow based?
BenchFlow is headquartered in San Francisco, United States, in the North America region.
How does BenchFlow make money?
One revenue line is on record: research/Open Source.
Who are BenchFlow's main competitors?
Direct peers on record are Braintrust, LangChain (LangSmith), Patronus AI and Galileo (Galileo AI). Emerging players are Anyscale and Confident AI (DeepEval). Broad incumbents are Weights & Biases, Hugging Face, OpenAI (Evals framework) and Anthropic (Claude Evals).
Does BenchFlow have an API?
No public API is recorded for BenchFlow.
What industry is BenchFlow in?
BenchFlow's product category is AI Agent Evaluation Software. Its primary akta.pro industry code is HDAAAAAJ, AI Benchmarking, Profiling & Performance Monitoring, with a secondary code of HDAAACAF, Agents & Autonomous Workflows (Tool Use, Planning, Multi-Agent).