Common Crawl
Common Crawl Foundation is a 501(c)(3) nonprofit that maintains a free, open repository of 300+ billion web pages spanning 15 years, used by AI labs including OpenAI, academic researchers, and web scientists for training language models and conducting research.
- Company typePrivate
- Founded2007
- HeadquartersBeverly Hills, United States
- Headcount1–10
- GTM typeB2B
- OfferingSoftware
What Common Crawl does
Common Crawl Foundation is a California 501(c)(3) nonprofit founded in 2007 by Gil Elbaz, operating from Beverly Hills, CA, with a lean 1-10 person engineering team. The organization maintains a free, open repository of web crawl data, currently containing over 300 billion pages spanning 15 years of internet history and adding 3-5 billion new pages each month through its proprietary ccBot crawler. Core products include the Web Crawl Data Repository (distributed in WARC, WAT, and WET formats), Web Graphs at the host and domain level (latest release covering April-June 2026 with 247.3M host nodes / 6.3B edges and 121.1M domain nodes / 3.9B edges), a CDXJ/URL Index with API access, and a recently launched AI Agent for natural-language search. The dataset is cited in over 10,000 research papers and is used by AI companies including OpenAI for training large language models, alongside academic and web-science research.
The technical stack centers on a distributed crawling infrastructure (ccBot) with data processing pipelines built in Python, Spark/PySpark, and the WebGraph framework, producing outputs in standardized web-archive formats (WARC, WAT, WET) plus a ZipNum-based CDX index. Distribution is global and entirely free, delivered via AWS S3 (under Amazon's Open Data Sponsorship Program, which eliminates the petabyte-scale hosting cost burden) and HTTPS endpoints, with an additional distribution surface on Hugging Face for ML practitioners.
The business model is nonprofit donation- and sponsorship-funded rather than transactional: there is no pricing, no paid tier, and no direct revenue from data licensing. AWS provides infrastructure at no cost through its Open Data Sponsorship Program. Operating geographies are global with no restrictions, and customer segments split between academic researchers (primary persona), AI/ML companies (primary business unit, including OpenAI and various AI labs), and web accessibility researchers. The organization faces ongoing regulatory and reputational pressure from major news publishers (CNN, NBC, USA Today, News/Media Alliance) regarding the inclusion of paywalled content in the archive, which it has begun addressing through an Opt-Out Registry.
Common Crawl firmographics
Firmographics- Name
- Common Crawl
- Legal name
- Common Crawl Foundation
- Website
- https://commoncrawl.org
- Company type
- Private
- Founded year
- 2007
- Operating status
- Operating
- Headcount range
- 1–10 employees
- Short description
- Common Crawl Foundation is a 501(c)(3) nonprofit that maintains a free, open repository of 300+ billion web pages spanning 15 years, used by AI labs including OpenAI, academic researchers, and web scientists for training language models and conducting research.
- Ownership category
- akta.pro rank
Common Crawl industry classification
Industry- Product category
- Open Web Data Infrastructure
- NAICS
- Web Search Portals and All Other Information Services (51929), Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (5182)
- SIC
- Services-Computer Programming, Data Processing, Etc. (7370)
- akta.pro primary industry
- Web Archiving, Cache & Snapshot Services (BPAMAAAL)
- akta.pro secondary industries
- Enterprise Search, Indexing & Content Discovery (HDAEAGAH), News & Content Aggregation Portals (BPAMAAAI)
Keywords
Where Common Crawl is headquartered
LocationHeadquarters
- HQ city
- Beverly Hills
- HQ country
- United States
- HQ region
- North America
Offices1 record
Markets served
Common Crawl business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Infrastructure, Technology or R&D, Personnel, Operations
Revenue model
- Donations and Sponsorships: As a 501(c)(3) nonprofit, Common Crawl relies on donations from individuals and corporate sponsors to fund operations. The organization actively seeks corporate sponsors to partner with for their non-profit open data mission.
- AWS Open Data Sponsorship: Cloud hosting infrastructure is provided free of charge through Amazon Web Services' Open Data Sponsorship Program, significantly reducing operational costs
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Freemium | Pay-as-you-go | Free open access to all web crawl data |
Go-to-market motion1 record
Distribution channels3 records
Marketing channels5 records
Common Crawl product offering
Product offeringCore offering
Common Crawl operates a large-scale web crawler (ccBot) that systematically archives the open web and publishes the resulting data as a free, open repository. Its core offerings include bulk web crawl archives in WARC/WAT/WET formats, host-level and domain-level web graphs, CDXJ and URL indexes, and an Index API, all hosted on AWS S3 under the AWS Open Data Sponsorship Program and accessible to anyone without charge.
Product overview
Common Crawl is a 501(c)(3) nonprofit organization that maintains a free, open repository of web crawl data. The core offering consists of the Web Crawl Data Repository containing over 300 billion web pages, supplemented by Web Graphs (host-level and domain-level hyperlink graphs), CDXJ Index (capture index format), URL Index (for discovering URLs across crawls), and WARC/WAT/WET file formats for raw and processed crawl data. Additional offerings include CCBot (the crawler), AI Agent (natural language search), and integration with Hugging Face for dataset access. The organization publishes 3-5 billion new pages monthly and data is hosted on AWS S3 under their Open Data Sponsorship Program.
Differentiator
Problem solved
Functional benefit
Products and services
- Web Crawl Data Repository Free, open repository of web crawl data containing over 300 billion pages spanning 15 years, with 3-5 billion new pages added each month. Data is distributed in WARC, WAT, and WET formats on AWS S3 for researchers, academics, and AI companies.
- Web Graphs Host-level and domain-level web graphs derived from crawl data using the WebGraph framework, used for ranking analysis, link spam detection, and large-scale graph research.
- CDXJ Index Capture Index JSON (CDXJ) line-format index enabling efficient lookup of archived web content by URL and timestamp, using the ZipNum CDX format for billions of pages.
- URL Index Index service for searching and discovering URLs across all crawl archives, including a columnar URL index for enhanced querying.
- Index API (CDX Server API) Public API at index.commoncrawl.org for querying crawl data, supporting wildcard URL queries, pagination, and JSON output for efficient programmatic lookup of archived URLs.
Quantifiable outcome
- Over 10,000 research papers have cited Common Crawl data
- +1 more outcomes
Companies that use Common Crawl
Customer profileNamed customers3 records
Segments3 records
Ideal customer profiles3 records
Common Crawl technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
AI capability2 records
Feature4 records
Common Crawl partnerships and signals
Strategic signalPartnerships
Two partnerships are on record, tiered core.
- Amazon Web Services (AWS)coreAWS provides hosting for Common Crawl's data through their Open Data Sponsorship Program, allowing the nonprofit to store and distribute petabytes of web crawl data at no cost. This is critical infrastructure for the organization's mission.
- WebGraph FrameworkcoreCommon Crawl uses the WebGraph framework (from University of Milano) for computing web graph properties and rankings. This open-source software enables large-scale graph analysis of the crawled data.
Scale indicators6 records
Recent moves6 records
Expansion highlights5 records
Common Crawl competitors and assessment
Company assessmentBroad incumbents
- Diffbot: Diffbot is a commercial web-scale knowledge extraction company that builds structured knowledge graphs from crawled web data. It addresses the same underlying market of structured web knowledge used for AI and enterprise search, but via a paid product rather than free open data.
Direct peers
- Internet Archive: The Internet Archive's Wayback Machine is the most direct peer — a nonprofit operating large-scale web archiving infrastructure (WARC files, historical snapshots) freely accessible to the public. Both organizations preserve historical web content at scale using similar file formats and serve overlapping research and AI-training use cases.
- Hugging Face: Hugging Face hosts Common Crawl's datasets and serves as a primary distribution channel for ML practitioners consuming them. Both operate as open-data infrastructure for the AI community and share overlapping users (AI researchers and model builders).
- Wikimedia Foundation: Wikimedia Foundation operates a similar nonprofit, open-data model for Wikipedia and structured knowledge, frequently used as training data alongside Common Crawl. Both serve as foundational open-data infrastructure for AI and academic research with comparable governance and funding structures.
Others
- Common Sense Media (Common Sense): Common Sense Media is a nonprofit producing open datasets for AI safety, child development, and media evaluation. It shares the nonprofit open-data model for AI training and research, although the focus areas differ.
- AWS Open Data Sponsorship Program: AWS's Open Data Sponsorship Program hosts Common Crawl and a broader catalog of petabyte-scale open datasets (e.g., NASA, NIH genomic data). It is an ecosystem enabler rather than a competitor, but its terms directly shape Common Crawl's cost structure and distribution reach.
Emerging players
- Archive.today: Archive.today is a smaller web archiving service capturing on-demand snapshots of web pages. While it operates at far smaller scale than Common Crawl, it addresses the same archival use case and competes for the same publisher opt-out and copyright narrative around AI training data.
- Curlie / Open Directory Project successors: Curlie and similar community-curated web directories historically served as open indexes of web content for research and discovery. They share Common Crawl's nonprofit, community-driven ethos but operate at much smaller scale and with manual curation.
- Open Web Search (open-source web search initiative): The Open Web Search EU initiative is a European open-web-indexing project aiming to provide an open alternative to commercial search indexes. It is comparable to Common Crawl in mission (open web data) but operates as a research consortium rather than a continuously archived dataset.
- Pushshift (now defunct, archived datasets): Pushshift historically provided free open datasets of social media data for research, similar in mission to Common Crawl but focused on Reddit and social platforms. Although largely defunct, its archived datasets remain a comparable open-data precedent in AI research.
Market position
Strengths4 records
Weaknesses1 record
Competitive moat4 records
Key risks5 records
Key highlights6 records
Customer concentration
Common Crawl social profiles
Digital presenceCommon Crawl financial estimates
Financial estimateRevenue estimate
Valuation estimate
Common Crawl leadership team
Management profileNumber of profiles
Profiles5 records
Common Crawl funding detail
Funding detailFunding overview
Funding rounds
Investors
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Common Crawl M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Common Crawl
What does Common Crawl do?
Common Crawl operates a large-scale web crawler (ccBot) that systematically archives the open web and publishes the resulting data as a free, open repository. Its core offerings include bulk web crawl archives in WARC/WAT/WET formats, host-level and domain-level web graphs, CDXJ and URL indexes, and an Index API, all hosted on AWS S3 under the AWS Open Data Sponsorship Program and accessible to anyone without charge.
Is Common Crawl a public or private company?
Common Crawl is a private company. It is classified as nonprofit foundation owned and is currently operating.
When was Common Crawl founded?
Common Crawl was founded in 2007. It employs 1 to 10 people.
Where is Common Crawl based?
Common Crawl is headquartered in Beverly Hills, United States, in the North America region.
How does Common Crawl make money?
Two revenue lines are on record. Donations and Sponsorships are the primary driver. The others are AWS Open Data Sponsorship.
Who are Common Crawl's main competitors?
Diffbot is listed as a broad incumbent. Direct peers are Internet Archive, Hugging Face and Wikimedia Foundation. Others are Common Sense Media (Common Sense) and AWS Open Data Sponsorship Program. Emerging players are Archive.today, Curlie / Open Directory Project successors, Open Web Search (open-source web search initiative) and Pushshift (now defunct, archived datasets).
Does Common Crawl have an API?
Yes. Common Crawl provides public API access via the Index API at index.commoncrawl.org for querying crawl data. The CDX Server API allows querying URLs from specific crawls, with support for wildcard queries, pagination, and JSON output. They also offer an open data repository accessible via AWS S3 and HTTPS endpoints. Developer documentation is at index.commoncrawl.org.
What industry is Common Crawl in?
Common Crawl's product category is Open Web Data Infrastructure. Its primary akta.pro industry code is BPAMAAAL, Web Archiving, Cache & Snapshot Services, with a secondary code of HDAEAGAH, Enterprise Search, Indexing & Content Discovery. Its NAICS code is 51929 and its SIC code is 7370.