MLCommons
MLCommons Association is a non-profit consortium developing open MLPerf benchmarks and AILuminate AI safety benchmarks, serving 125+ member organizations across chip vendors, hyperscalers, AI labs, and academia.
- Company typePrivate
- Founded2018
- HeadquartersSan Francisco, United States
- Headcount1–10
- GTM typeB2B
- OfferingSoftware
What MLCommons does
MLCommons Association is a US-based non-profit consortium founded in 2018 that develops open, industry-standard benchmarks and data tooling for measuring AI quality, performance, and safety. Its flagship offerings span the MLPerf benchmark family — Training, Inference (Datacenter, Edge, Mobile, Tiny), Storage, Client, Automotive, and Endpoints — plus the AILuminate safety suite (twelve hazard categories, 24,000+ test prompts per language), AlgoPerf training-algorithm benchmarks, the Croissant metadata standard (700k+ datasets), MLCube containerization, and Chakra execution traces.
The organization's technical proposition is that of a neutral, peer-reviewed measurement infrastructure: MLPerf submissions are bi-annual and results are publicly published, with hardware vendors (NVIDIA Blackwell Ultra, AMD GPUs, Intel Xeon 6, Qualcomm Snapdragon), hyperscalers (Microsoft Azure, Oracle, Google), and AI labs (Meta, Shopify-side collaborations) competing on leaderboards. AILuminate and the AI Risk & Reliability working group extend the measurement layer into agentic, multimodal/multilingual, and security benchmarks, increasingly aligned with emerging regulation (EU AI Act, Colorado AI Act) and government-led third-party evaluations (UK AI Security Institute, U.S. NSF NAIRR pilot).
MLCommons operates a membership-based revenue model with 125+ member organizations contributing membership fees (undisclosed publicly); membership is required for participation in most working groups, while non-member submissions are permitted in select public groups under a Non-member Test Agreement. Distribution is primarily digital — GitHub repositories, a public results dashboard at mlcommons.org/visualizer, mobile apps via Apple App Store and Google Play, and benchmark demonstrations at industry venues such as GTC 2026. Leadership is anchored by President Peter Mattson (former Google/NVIDIA research) with working-group chairs drawn from Google DeepMind, Qualcomm, Microsoft CoreAI, AMD, NVIDIA, Intel, and leading academic institutions.
MLCommons firmographics
Firmographics- Name
- MLCommons
- Legal name
- MLCommons Association
- Website
- https://mlcommons.org
- Company type
- Private
- Founded year
- 2018
- Operating status
- Operating
- Headcount range
- 1–10 employees
- Short description
- MLCommons Association is a non-profit consortium developing open MLPerf benchmarks and AILuminate AI safety benchmarks, serving 125+ member organizations across chip vendors, hyperscalers, AI labs, and academia.
- Ownership category
- akta.pro rank
MLCommons industry classification
Industry- Product category
- AI Benchmarking and Standards
- NAICS
- Professional Organizations (813920)
- SIC
- Services-Testing Laboratories (8734)
- akta.pro primary industry
- End-to-End MLOps & ML Platform Suites (HDAAABAA)
- akta.pro secondary industries
- Model Testing, Validation & Quality Assurance (HDAAABAH), ML Lifecycle Collaboration & Workspace Management (HDAAABAK)
Keywords
Where MLCommons is headquartered
LocationHeadquarters
- HQ city
- San Francisco
- HQ country
- United States
- HQ region
- North America
Markets served
MLCommons business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Personnel, Technology or R&D, Operations, Others
Revenue model
- Membership Fees: MLCommons operates as a membership-based consortium with over 125 members including startups, leading companies, academics, and non-profits. Membership is required for participation in most benchmark working groups and provides access to collaboration resources and decision-making processes.
Go-to-market motion1 record
Distribution channels5 records
Marketing channels8 records
MLCommons product offering
Product offeringCore offering
MLCommons is a non-profit consortium that develops open industry-standard benchmarks and data tooling for measuring AI system quality, performance, and risk. Its core offerings include the MLPerf benchmark suites (Training, Inference Datacenter/Edge/Mobile/Tiny, Storage, Client, Automotive, Endpoints), AILuminate AI safety benchmarks covering 12 hazard categories, AlgoPerf training algorithm benchmarks, and the Croissant metadata standard used by 700k+ datasets.
Product overview
MLCommons is a nonprofit organization operating as a benchmarking and standards body rather than a traditional product company. Its portfolio consists of industry-standard benchmark suites, data standards, and research tools. The core MLPerf family includes MLPerf Training (measuring model training speed), MLPerf Inference variants for Datacenter, Edge, Mobile, and Tiny environments, plus specialized suites for Automotive and Client systems. AILuminate provides AI safety benchmarking across twelve hazard categories. AlgoPerf measures training algorithm optimization speedups. Supporting infrastructure includes MLCube containers, the Croissant metadata standard (used by 700,000+ datasets), and Chakra execution traces. The organization publishes results from 89.7K+ benchmark submissions and maintains over 125 member organizations across industry and academia.
Differentiator
Problem solved
Functional benefit
Brands
- AILuminate: AI safety and security benchmarks assessing genAI across 12 hazard categories.
- AlgoPerf
- MLPerf Training
- MLPerf Inference
- MLPerf Mobile
- MLPerf Client
- MLPerf Automotive
- MLPerf Inference Tiny
- MLPerf Storage
- AILuminate Jailbreak
Products and services
- MLPerf Training Industry-standard benchmark measuring how fast systems can train AI models to a target quality metric, covering LLMs, image generation, recommendation, GNNs, BERT, object detection, and speech recognition.
- MLPerf Inference: Datacenter Benchmark measuring how fast datacenter systems process inputs and produce results using trained models across Server, Offline, Single Stream, and Multiple Stream scenarios with latency constraints.
- MLPerf Inference: Edge Benchmark measuring AI inference performance on edge systems, testing how fast systems can process inputs and produce results using trained models.
- MLPerf Inference: Mobile Benchmark measuring AI performance on consumer mobile devices (smartphones, tablets, notebooks), testing LLMs, image classification, object detection, segmentation, and language processing.
- MLPerf Inference: Tiny Benchmark measuring AI inference performance on ultra-low-power systems including image classification, keyword spotting, person detection, and anomaly detection.
- MLPerf Storage Benchmark measuring storage system performance for AI training workloads, testing how fast storage can supply training data to keep GPUs/accelerators fed.
- MLPerf Client Benchmark evaluating AI performance on personal computers (laptops, desktops, workstations), testing LLMs such as Llama 3.1 8B and Phi models for code analysis, content generation, and summarization tasks.
- MLPerf Automotive Benchmark suite measuring AI performance for automotive ADAS/AD perception and In-Vehicle Infotainment systems, testing 2D/3D object detection and semantic segmentation on Cognata and nuScenes datasets.
- MLPerf Endpoints Benchmark for evaluating GenAI API endpoint performance using an OpenAI-compatible API interface, with Pareto curves measuring throughput vs interactivity tradeoffs across the full operating range.
- AILuminate AI safety and security benchmark suite assessing generative AI systems across twelve hazard categories, including Safety v1.0 for safety response evaluation and Jailbreak v0.5 for adversarial prompt resistance testing.
- AlgoPerf: Training Algorithms Benchmark measuring neural network training speedups due to algorithmic improvements (optimizer, hyperparameters), using eight fixed workloads on standardized 8x V100 GPU hardware.
- Croissant Metadata Standard Metadata format standardizing how ML datasets are described, making ML work easier to reproduce and replicate, adopted by over 700,000 datasets.
- MLCube Containerization standard for packaging and running ML models consistently, improving AI ease-of-use and enabling AI to scale to more users.
Quantifiable outcome
- 89.7k+ MLPerf Performance Results to-date
- +3 more outcomes
Companies that use MLCommons
Customer profileNamed customers11 records
Segments5 records
Ideal customer profiles4 records
MLCommons technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- No
- API docs
- API detail
Core technology
AI maturity
App detail
AI capability15 records
Feature8 records
MLCommons partnerships and signals
Strategic signalPartnerships
18 partnerships are on record, tiered minor, core and major.
- New York Academy of SciencesminorCo-hosted the 3rd annual 'New Wave of AI in Healthcare' symposium with Icahn School of Medicine at Mount Sinai, discussing AI's expanding role in healthcare delivery, diagnostics, and drug discovery.
- Icahn School of Medicine at Mount SinaiminorCo-hosted the 3rd annual 'New Wave of AI in Healthcare' symposium with the New York Academy of Sciences, bringing together researchers, clinicians, and industry leaders to discuss AI's role in healthcare.
- Autonomous Vehicle Compute Consortium (AVCC)coreMLCommons partners with AVCC on the MLPerf Automotive benchmark suite. AVCC provides technical reports (TR003, TR004, TR007) and MLCommons provides the benchmarking infrastructure. The collaboration focuses on developing industry-standard ML benchmarks for automotive computing platforms used in ADAS/AD and IVI systems.
- AMDcoreAMD actively submits results to MLPerf benchmark suites, including Inference v6.0 and Training v6.0. AMD showcases its GPU performance through MLCommons benchmarks as part of its AI hardware validation.
- NVIDIAcoreNVIDIA is a key benchmark submitter to MLPerf suites. NVIDIA's Blackwell Ultra architecture topped the MLPerf Inference v5.1 leaderboard, achieving up to 5,842 tokens per second per GPU offline.
- GooglecoreGoogle submits benchmark results and contributes to working groups. Google DeepMind team members lead multimodal safety benchmarking workstream within the AI Risk & Reliability working group.
- Microsoft AzuremajorMicrosoft Azure submits cloud-based AI training benchmark results and contributes to MLPerf Training working groups. Microsoft CoreAI team works on AI evaluation tooling.
- OraclemajorOracle participated in MLPerf Training v6.0 benchmark submissions as one of 24 organizations submitting results.
- MetacoreMeta collaborates with MLCommons on MLPerf Inference benchmarks and is listed as a key collaboration partner for increased adoption of large-scale multi-node systems.
- ShopifymajorShopify collaborates with MLCommons on MLPerf Inference benchmarks, contributing to the diverse set of industry participants.
- UltralyticsmajorUltralytics collaborates on MLPerf Inference benchmarks as part of the expanding ecosystem of AI system developers.
- IBMmajorIBM Storage Scale participates in MLPerf Storage v2.0 benchmarks, achieving 656.7 GiB/s read bandwidth for 1T parameter model training.
- Singapore Ministry of Digital Development and InformationmajorGoogle and Singapore's Ministry expanded their National AI Partnership, building on a 2022 agreement to deploy AI across public services. This partnership will serve as a model for national AI programs.
- U.S. National Science Foundation (NSF)majorNSF launched the National AI Research Resource (NAIRR) pilot, partnering with MLCommons and 25 other organizations to democratize AI research access. Initial focus on safe, secure AI and healthcare applications.
- UK AI Security InstitutecoreThe layered assurance ecosystem combines lab self-attestation with government institute validation from UK AI Security Institute and independent third-party evaluation from organizations including MLCommons.
- Oxford Internet InstituteminorResearchers from Oxford Internet Institute and Nuffield Department of Primary Care Health Sciences at the University of Oxford collaborated with MLCommons on a medical advice study comparing AI chatbots against traditional methods.
- Nuffield Department of Primary Care Health Sciences, University of OxfordminorCollaborated with MLCommons on a study involving 1,298 people in the UK comparing effectiveness of AI chatbots (GPT-4o, Llama 3, Command R) against traditional methods for assessing medical symptoms.
- CognataminorCognata provides the MLCommons Cognata Dataset used in MLPerf Automotive benchmarks for 2D object detection and semantic segmentation testing.
Scale indicators6 records
Recent moves6 records
Expansion highlights7 records
MLCommons competitors and assessment
Company assessmentDirect peers
- HELM (Stanford CRFM): Stanford's Holistic Evaluation of Language Models is a leading academic benchmark for evaluating large language models across accuracy, calibration, robustness, fairness, and efficiency. Directly comparable to MLCommons as a primary AI benchmarking initiative targeting transparency and reproducibility.
- LMSYS Chatbot Arena: Large Model Systems Organization's crowdsourced LLM benchmark using pairwise human preference votes. Comparable to MLCommons as a widely-cited third-party AI evaluation platform with community-driven methodology.
- Hugging Face Open LLM Leaderboard: Hugging Face's benchmark for open-source LLMs evaluating reasoning, knowledge, and truthfulness. Comparable to MLCommons' Inference benchmarks as a public-facing AI evaluation reference for the open-source community.
- Partnership on AI: Multi-stakeholder nonprofit organization developing best practices on AI safety, fairness, and governance. Directly comparable to MLCommons as a similar consortium model with industry, academia, and civil society members focused on responsible AI development.
Broad incumbents
- Apache Software Foundation: Open-source software foundation operating as a membership-driven consortium with vendor-neutral governance. Comparable to MLCommons as a precedent for industry consortium structures that develop shared technical standards with member-funded operations.
- Linux Foundation: Largest open-source foundation supporting Linux, Kubernetes, and other critical infrastructure projects. Comparable to MLCommons as a mature model of vendor-neutral consortium governance and standards stewardship that MLCommons aspires to for AI benchmarking infrastructure.
- IEEE Standards Association: Major international standards body developing technical standards across industries including AI. Comparable to MLCommons as a precedent for member-driven technology standardization, though much larger and more established.
Emerging players
- OpenAI Evals: OpenAI's open-source evaluation framework for testing LLM capabilities. Comparable as an emerging benchmark initiative but single-vendor-driven, contrasting with MLCommons' multi-stakeholder approach.
- EleutherAI Evaluation Harness: Open-source framework for evaluating autoregressive language models across hundreds of standardized tasks. Comparable to MLCommons as a community-driven AI evaluation tool, though focused primarily on academic and open-model use cases.
Others
- MLPerf (benchmark suite brand): Listed here for completeness only — MLPerf is MLCommons' flagship product and trademark, not a separate peer.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat5 records
Key risks6 records
Key highlights6 records
Customer concentration
MLCommons social profiles
Digital presenceMLCommons financial estimates
Financial estimateRevenue estimate
Valuation estimate
MLCommons leadership team
Management profileNumber of profiles
Profiles8 records
MLCommons funding detail
Funding detailFunding overview
Funding rounds
Investors
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
MLCommons M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about MLCommons
What does MLCommons do?
MLCommons is a non-profit consortium that develops open industry-standard benchmarks and data tooling for measuring AI system quality, performance, and risk. Its core offerings include the MLPerf benchmark suites (Training, Inference Datacenter/Edge/Mobile/Tiny, Storage, Client, Automotive, Endpoints), AILuminate AI safety benchmarks covering 12 hazard categories, AlgoPerf training algorithm benchmarks, and the Croissant metadata standard used by 700k+ datasets.
Is MLCommons a public or private company?
MLCommons is a private company. It is classified as nonprofit foundation owned and is currently operating.
When was MLCommons founded?
MLCommons was founded in 2018. It employs 1 to 10 people.
Where is MLCommons based?
MLCommons is headquartered in San Francisco, United States, in the North America region.
How does MLCommons make money?
One revenue line is on record: membership Fees.
Who are MLCommons's main competitors?
Direct peers on record are HELM (Stanford CRFM), LMSYS Chatbot Arena, Hugging Face Open LLM Leaderboard and Partnership on AI. Broad incumbents are Apache Software Foundation, Linux Foundation and IEEE Standards Association. Emerging players are OpenAI Evals and EleutherAI Evaluation Harness. MLPerf (benchmark suite brand) is listed as an others.
Does MLCommons have an API?
No public API is recorded for MLCommons.
What industry is MLCommons in?
MLCommons's product category is AI Benchmarking and Standards. Its primary akta.pro industry code is HDAAABAA, End-to-End MLOps & ML Platform Suites, with a secondary code of HDAAABAH, Model Testing, Validation & Quality Assurance. Its NAICS code is 813920 and its SIC code is 8734.