HTTP Archive
HTTP Archive is a community-run non-profit that has tracked web technology adoption and performance since 2010 by crawling over 16 million sites monthly and publishing the data and annual Web Almanac reports free via Google BigQuery for developers, researchers, and vendors.
- Company typePrivate
- Founded2010
- HeadquartersSan Francisco, United States
- Headcount1–10
- GTM typeB2B
- OfferingSoftware
What HTTP Archive does
HTTP Archive is a community-run non-profit project founded in 2010 in San Francisco that systematically measures how the public web is built and how it performs. Its core product is a web crawling infrastructure that uses WebPageTest and Lighthouse to test metadata for over 16 million websites every month, capturing request/response data, response bodies, execution traces, and web platform API usage. This raw crawl corpus is published as a free public dataset on Google BigQuery, enabling longitudinal and cross-sectional analysis of web technology adoption and performance.
The organization translates this raw measurement data into human-readable insight through the annual Web Almanac, which in its 2025 edition spans 16 chapters covering fonts, WebAssembly, third parties, generative AI, SEO, accessibility, performance, privacy, security, capabilities, PWA, CMS, ecommerce, page weight, CDN, and cookies. A supplementary product, the Core Web Vitals Technology Report, joins HTTP Archive's origin-level technology identification with Chrome UX Report performance data to link technology choices with real-world user-experience outcomes. Local editions of the Almanac are produced in 12+ languages, and all reports are distributed freely as web pages and downloadable PDFs. A community discussion forum and open GitHub repositories support the contributor and analyst ecosystem.
HTTP Archive has no commercial revenue model: the dataset is free, reports are free, and there are no pricing tiers or paid products. Operations are sustained by a small (1-10) core team, 74 named volunteer contributors for the 2025 Web Almanac, in-kind infrastructure from Google (BigQuery and Chrome UX Report), and a reciprocal relationship with the Internet Archive. Named administrators include Rick Viscomi and Patrick Meenan, with community moderators such as Max Ostapenko and Barry Pollard. Key technology partners are Google (BigQuery hosting and CrUX data) and the Internet Archive.
HTTP Archive firmographics
Firmographics- Name
- HTTP Archive
- Legal name
- HTTP Archive
- Website
- https://httparchive.org
- Company type
- Private
- Founded year
- 2010
- Operating status
- Operating
- Headcount range
- 1–10 employees
- Short description
- HTTP Archive is a community-run non-profit that has tracked web technology adoption and performance since 2010 by crawling over 16 million sites monthly and publishing the data and annual Web Almanac reports free via Google BigQuery for developers, researchers, and vendors.
- Ownership category
- akta.pro rank
HTTP Archive industry classification
Industry- Product category
- Web Performance Analytics
- NAICS
- Web Search Portals and All Other Information Services (51929), Web Search Portals and All Other Information Services (519290), Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (51821)
- SIC
- Services-Computer Processing & Data Preparation (7374)
- akta.pro primary industry
- Technical SEO (Audits, Crawl/Indexing, Site Architecture) (BPAFAJAA)
- akta.pro secondary industries
- Meta-Search & Search Aggregators (BPAMAAAC), News & Content Aggregation Portals (BPAMAAAI)
Keywords
Where HTTP Archive is headquartered
LocationHeadquarters
- HQ city
- San Francisco
- HQ country
- United States
- HQ region
- North America
Markets served
HTTP Archive business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Technology or R&D, Infrastructure, Personnel, Operations
Revenue model
- Public Good / Open Data Service: HTTP Archive is a community-run project that provides free access to web performance data. No commercial revenue model exists; the project operates as a non-profit open-source initiative supported by community contributions and likely Google infrastructure donations.
Go-to-market motion1 record
Distribution channels3 records
Marketing channels6 records
HTTP Archive product offering
Product offeringCore offering
HTTP Archive periodically crawls the top sites on the web using WebPageTest and Lighthouse, recording detailed metadata about fetched resources, web platform APIs, and execution traces for over 16 million websites each month. The resulting dataset is made freely available via Google BigQuery as a public dataset, and feeds published reports including the annual Web Almanac, Core Web Vitals Technology Report, State of the Web trends, and Project Fugu capabilities tracking.
Product overview
HTTP Archive is a community-run project that tracks how the web is built. The core offering consists of a web crawling infrastructure that periodically tests over 16 million websites using WebPageTest and Lighthouse, storing detailed metadata about resources, APIs, and execution traces. This data powers multiple products: the Web Almanac (annual comprehensive reports), the Core Web Vitals Technology Report (performance tracking), and State of the Web Reports (trends in network utilization and web standards). Data is accessible via Google BigQuery as a public dataset. The Capabilities Report tracks Project Fugu web platform API adoption. A community discussion forum supports user engagement and analysis sharing.
Differentiator
Problem solved
Functional benefit
Brands
- Web Almanac: Annual state of the web report combining raw stats and trends with web community expertise
- Core Web Vitals Tech Report
Products and services
- Web Almanac HTTP Archive's annual comprehensive report on the state of the web, combining raw statistics and trends with expertise from the web community. Available in 12+ languages with downloadable PDF ebook.
- Core Web Vitals Technology Report Interactive tool for tracking and comparing web technology performance and adoption, providing real-world Core Web Vitals performance data from the Chrome UX Report at the origin level with technologies identified by HTTP Archive.
- State of the Web Reports Long-term tracking reports of web metrics including total kilobytes, total requests, TCP connections per page, adoption of techniques for efficient network utilization, and usage of web standards like HTTPS.
- Public Dataset (Google BigQuery) Free public dataset hosted on Google BigQuery providing detailed archived information about crawled websites, including request/response metadata, response bodies, and execution traces. Users can run SQL queries or download data for offline analysis.
- Capabilities Report (Project Fugu) Tracks the adoption of web platform capabilities (Project Fugu) including APIs such as Async Clipboard, Badging API, and Screen Wake Lock that enable web apps to match native app functionality.
Quantifiable outcome
- 591% increase in WebGPU adoption highlighted in 2025 Generative AI chapter
- +2 more outcomes
Companies that use HTTP Archive
Customer profileSegments3 records
Ideal customer profiles3 records
HTTP Archive technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- No
- API docs
- API detail
Core technology
AI maturity
App detail
Integration1 record
Feature3 records
HTTP Archive partnerships and signals
Strategic signalPartnerships
Three partnerships are on record, tiered core.
- Internet ArchivecoreHTTP Archive collaborates with Internet Archive, displaying their logo on the website. Internet Archive provides archiving infrastructure that complements HTTP Archive's web measurement mission.
- Google (BigQuery, Chrome UX Report)coreHTTP Archive data is hosted as a public dataset on Google BigQuery. Core Web Vitals Tech Report integrates real-world performance data from Google's Chrome UX Report at the origin level. Uses Google infrastructure for data processing.
- Web Almanac Chapter Contributorscore74 volunteer contributors from the web community contributed to the 2025 Web Almanac across planning, research, writing, and production phases.
Scale indicators5 records
Recent moves6 records
Expansion highlights5 records
HTTP Archive competitors and assessment
Company assessmentDirect peers
- Wappalyzer: Wappalyzer identifies the technologies used on websites (CMS, frameworks, analytics, CDNs) at scale — overlapping directly with HTTP Archive's website technology detection and adoption metrics.
- BuiltWith: BuiltWith tracks web technology adoption across millions of websites, providing technology profiling, lead generation, and trend data — directly comparable to HTTP Archive's technology adoption and trend tracking use case.
- W3Techs: W3Techs publishes statistics on web technology usage (CMS, server-side languages, JavaScript libraries) by surveying websites — comparable longitudinal technology adoption tracking to HTTP Archive.
- StatCounter Global Stats: StatCounter tracks web technology adoption (browsers, operating systems, search engines) through aggregated usage data, similar to HTTP Archive's adoption trend reporting across the web ecosystem.
Broad incumbents
- Google Lighthouse: Lighthouse is the auditing tool powering HTTP Archive's website evaluations, offered by Google as a standalone product for developers to run performance, SEO, accessibility, and best-practice audits on their own sites.
- PageSpeed Insights (Google): Google's PageSpeed Insights combines Lighthouse lab data with CrUX field data to score website performance — overlapping with HTTP Archive's Core Web Vitals Tech Report for individual-site performance evaluation.
- WebPageTest (Catchpoint): WebPageTest is the underlying synthetic testing tool HTTP Archive uses; operated by Catchpoint, it provides professional-grade performance testing with a commercial offering that overlaps HTTP Archive's measurement foundation.
- Chrome User Experience Report (CrUX): Google's CrUX provides real-user Core Web Vitals data that HTTP Archive integrates into its Core Web Vitals Tech Report; CrUX is the upstream authoritative source for the same metrics HTTP Archive enriches.
Emerging players
- Pingdom / GTmetrix: GTmetrix and Pingdom provide website performance and page weight analysis, surfacing metrics comparable to HTTP Archive's State of the Web and Page Weight reports at the individual-site level.
- Datanyze (now Crunchbase): Datanyze was a web technology tracking and lead generation platform acquired by Crunchbase; it maintained technology adoption intelligence comparable to HTTP Archive's coverage of web technology market share.
Market position
Strengths4 records
Weaknesses4 records
Competitive moat4 records
Key risks5 records
Key highlights6 records
Customer concentration
HTTP Archive social profiles
Digital presenceHTTP Archive financial estimates
Financial estimateRevenue estimate
Valuation estimate
HTTP Archive leadership team
Management profileNumber of profiles
Profiles4 records
HTTP Archive funding detail
Funding detailFunding overview
Funding rounds
Investors
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
HTTP Archive M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about HTTP Archive
What does HTTP Archive do?
HTTP Archive periodically crawls the top sites on the web using WebPageTest and Lighthouse, recording detailed metadata about fetched resources, web platform APIs, and execution traces for over 16 million websites each month. The resulting dataset is made freely available via Google BigQuery as a public dataset, and feeds published reports including the annual Web Almanac, Core Web Vitals Technology Report, State of the Web trends, and Project Fugu capabilities tracking.
Is HTTP Archive a public or private company?
HTTP Archive is a private company. It is classified as nonprofit foundation owned and is currently operating.
When was HTTP Archive founded?
HTTP Archive was founded in 2010. It employs 1 to 10 people.
Where is HTTP Archive based?
HTTP Archive is headquartered in San Francisco, United States, in the North America region.
How does HTTP Archive make money?
One revenue line is on record: public Good / Open Data Service.
Who are HTTP Archive's main competitors?
Direct peers on record are Wappalyzer, BuiltWith, W3Techs and StatCounter Global Stats. Broad incumbents are Google Lighthouse, PageSpeed Insights (Google), WebPageTest (Catchpoint) and Chrome User Experience Report (CrUX). Emerging players are Pingdom / GTmetrix and Datanyze (now Crunchbase).
Does HTTP Archive have an API?
No public API is recorded for HTTP Archive.
What industry is HTTP Archive in?
HTTP Archive's product category is Web Performance Analytics. Its primary akta.pro industry code is BPAFAJAA, Technical SEO (Audits, Crawl/Indexing, Site Architecture), with a secondary code of BPAMAAAC, Meta-Search & Search Aggregators. Its NAICS code is 51929 and its SIC code is 7374.