Internet Archive
The Internet Archive is a 501(c)(3) nonprofit founded in 1996 that operates the Wayback Machine and related free public archival services, preserving over one trillion web pages and 210+ petabytes of cultural, governmental, and media content for global public access.
- Company typePrivate
- Founded1996
- HeadquartersSan Francisco, United States
- Headcount51–100
- GTM typeB2C
- OfferingDigital Commerce or Content
What Internet Archive does
The Internet Archive is a 501(c)(3) nonprofit digital library founded in 1996 by Brewster Kahle and headquartered in San Francisco, California. The organization operates the Wayback Machine and a constellation of related free public archival services built on a custom, horizontally-scaled preservation platform spanning 210+ petabytes of web captures, digitized books, audio, video, television news, and software. As of October 2025, the Wayback Machine had preserved over one trillion web pages; the organization employs 51-100 staff and serves a global base of researchers, journalists, librarians, historians, and the general public.
The core technical architecture combines commodity distributed storage, scheduled and event-driven web crawling, and full-text search indexes (including the TV News Archive closed-captioning search). Product surface includes the Wayback Machine, Archive-It (an institutional subscription service for organizational web archiving), the Open Library for digitized books, and curated collections such as Democracy's Library and the Great 78 Project. Capture is increasingly integrated into publishing workflows through the February 2025 Automattic/WordPress partnership, which embeds one-click Save Page Now directly in the WordPress editor.
Revenue is generated through a diversified mix of individual donations averaging approximately $14 per gift, philanthropic grants from foundations such as Press Forward and the Filecoin Foundation, and institutional Archive-It subscriptions. Annual revenue is approximately $37 million. The organization is structurally constrained by ongoing publisher blocking (340+ news outlets, with an 87% drop in news-website archiving), copyright litigation exposure (Hachette v. IA ruling, Martino suit, Great 78 settlement), and the operational aftermath of the October 2024 cyberattack that compromised approximately 31 million user accounts.
Internet Archive firmographics
Firmographics- Name
- Internet Archive
- Legal name
- Internet Archive
- Website
- https://archive.org
- Company type
- Private
- Founded year
- 1996
- Operating status
- Operating
- Headcount range
- 51–100 employees
- Short description
- The Internet Archive is a 501(c)(3) nonprofit founded in 1996 that operates the Wayback Machine and related free public archival services, preserving over one trillion web pages and 210+ petabytes of cultural, governmental, and media content for global public access.
- Ownership category
- akta.pro rank
Internet Archive industry classification
Industry- Product category
- Digital Library and Web Archiving
- NAICS
- Libraries and Archives (51921), Libraries and Archives (519210), Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (5182)
- SIC
- Services-Computer Programming, Data Processing, Etc. (7370)
- akta.pro primary industry
- Digital Archives, Libraries & Institutional Access Platforms (MPAJAMAJ)
- akta.pro secondary industries
- Cloud Media Archive & Cold Storage Services (MPAHAGAF), Archive Restoration & Digitization (Film/Video/Audio Scanning, Remastering) (MPAHAGAI), Libraries, Archives & Special Collections (BPAGAHAF)
Keywords
Where Internet Archive is headquartered
LocationHeadquarters
- HQ city
- San Francisco
- HQ country
- United States
- HQ region
- North America
Offices1 record
Markets served
Internet Archive business model
Business model- GTM type
- B2C
- Offering type
- Digital Commerce or Content
- Cost components
- Infrastructure, Personnel, Technology or R&D, Operations, Others
Revenue model
- Donations: Primary revenue source. The organization is powered by online donations averaging about $14. Individual donors contribute to support operations and preservation efforts.
- Grants: Foundation and organizational grants support specific initiatives. Past funders include Kahle/Austin Foundation, Arcadia Foundation, John S. and James L. Knight Foundation, Sunlight Foundation, and others.
- Archive-It Subscription: Subscription service for institutions requiring curated web archiving solutions, professional features, and organizational support for web preservation projects.
- Merchandise: Sculptural recognition items (staff members who work at the organization for three years receive sculptures of themselves seated in the church pews) represent a unique non-revenue but community-building initiative.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Freemium | Pay-as-you-go | Free Access (Wayback Machine, Archive Browse) |
| Subscription | Annual | Archive-It Institutional Subscription |
Go-to-market motion3 records
Distribution channels6 records
Marketing channels5 records
Internet Archive product offering
Product offeringCore offering
The Internet Archive is a nonprofit digital library that preserves and provides free public access to a vast collection of digital content including web pages (via the Wayback Machine), books, audio recordings, video, software, and images. It operates one of the world's largest digital preservation infrastructures, storing over 210 petabytes of data and adding approximately 100 terabytes of new material daily. The organization offers free access to most services for individual users, while institutional customers can subscribe to Archive-It for curated web archiving solutions.
Product overview
The Internet Archive is a nonprofit digital library operating as a multi-product platform with the Wayback Machine as its flagship web archiving service at its core. The platform preserves over 866 billion web pages through the Wayback Machine, maintains over 210 petabytes of stored digital materials, and adds approximately 100 terabytes of new content daily. Beyond web archiving, the platform encompasses multiple specialized collections: the Open Library for books and ebook lending; the TV News Archive for searchable broadcast content; the Live Music Archive for concert recordings and the Grateful Dead collection; Internet Arcade and Console Living Room for browser-based retro gaming emulation; MS-DOS Games; 78 RPMs and Cylinder Recordings (Great 78 Project); and the Understanding 9/11 Television News Archive. Additional products include the Wayback Machine Link Fixer WordPress plugin (jointly developed with Automattic), Archive-It subscription service for institutional web archiving, mobile apps, and browser extensions. Recent initiatives include Democracy's Library for government data preservation and the 2026 Public Song Project with WNYC. The organization celebrated its 30th anniversary on May 10, 2026, and reached the milestone of preserving 1 trillion web pages in October 2025.
Differentiator
Problem solved
Functional benefit
Brands
- Wayback Machine: Digital archive service preserving copies of webpages since 1996, allowing users to revisit historical snapshots of websites.
- Open Library
- TV News Archive
Products and services
- Wayback Machine The flagship web archiving service that has been preserving copies of webpages since 1996, allowing users to revisit historical snapshots of websites from years or decades past. It has archived over 1 trillion web pages and serves as a critical tool for preserving the historical web record and combating link rot. Free public access.
- Open Library An online project to catalog every book ever published, providing free ebook lending, book metadata, and integration with the Internet Archive's digital library collections through the Open Library Explorer.
- Internet Archive TV News A research library service that repurposes closed captioning to enable users to search, quote, and borrow U.S. TV news programs. Contains more than 4.35 million news programs collected since 2009 from national U.S. networks and stations.
- Understanding 9/11 Television News Archive A library of news coverage of the events of September 11, 2001 and their aftermath as presented by U.S. and international broadcasters. Contains over 3,000 hours of TV news from 20 channels over 7 days for study, research, and analysis.
- Live Music Archive A collection of live music recordings, including the Grateful Dead collection and thousands of rare concert recordings from private collectors, preserved for free public streaming and download.
- Librivox A platform for free audiobooks, providing public domain audio recordings of books read by volunteers from around the world.
- Internet Arcade A collection of vintage arcade games preserved and made playable through browser-based emulation technology, allowing users to play classic games directly in the web browser.
- Console Living Room A collection of vintage console video games preserved and made playable through browser-based emulation.
- MS-DOS Games A collection of historical MS-DOS games preserved and made playable through browser-based emulation.
- 78 RPMs and Cylinder Recordings The Great 78 Project digitizes and distributes 78 RPM records and cylinder recordings, preserving historical sound recordings.
- Wayback Machine Link Fixer A WordPress plugin that automatically redirects users from dead links to archived versions of web pages, proactively archives content updates, and reverts links when original pages come back online. Developed in partnership with Automattic.
- Archive-It A subscription service that enables organizations to build and preserve collections of web content, providing tools for curating, searching, and managing archived web materials. Pricing not publicly disclosed; requires direct inquiry for institutional customers.
- Wayback Machine Mobile Apps Native mobile applications for iOS and Android that provide access to the Wayback Machine for archiving and retrieving web pages on mobile devices.
- Wayback Machine Browser Extensions Browser extensions for Chrome, Firefox, Safari, and Edge that provide direct Wayback Machine access and functionality from within the browser.
- Democracy's Library An initiative to preserve and make openly accessible government data and public records, including partnerships with organizations like the Government of Bermuda to upload public datasets to decentralized storage networks on the Filecoin network.
- NSA Clip Library A curated research library of TV news clips regarding the NSA, its oversight, and privacy issues from 2009-2014, with speaker indexing for researchers.
- Public Song Project A 2026 collaboration between WNYC and the Internet Archive inviting musicians to adapt, remix, or reimagine works from the public domain published in 1930 or earlier, with selected submissions featured by WNYC and preserved in the Internet Archive's digital library.
Quantifiable outcome
- 87% drop in news website archiving following publisher blocking actions (May-October 2025)
- +3 more outcomes
Companies that use Internet Archive
Customer profileNamed customers7 records
Segments6 records
Ideal customer profiles4 records
Internet Archive technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration5 records
AI capability7 records
Feature5 records
Internet Archive partnerships and signals
Strategic signalPartnerships
Five partnerships are on record, tiered core, major and supporting.
- Filecoin FoundationcoreCollaboration with Filecoin Foundation for the Democracy's Library initiative. The Government of Bermuda uploads public datasets to the Filecoin network through this partnership, with Internet Archive facilitating the initiative to enhance data resilience, transparency, and verifiability.
- Automattic (WordPress)coreStrategic partnership to integrate Wayback Machine functionality directly into WordPress infrastructure. The Wayback Machine Link Fixer plugin automatically archives external links in WordPress content and redirects users to archived versions when links break. WordPress powers more than 43% of all websites globally.
- Smithsonian Institution, Flickr Foundation, MIT Open Learning, Starling LabmajorMajor cultural and educational organizations uploading datasets to the Filecoin network for decentralized storage. The Internet Archive is part of this coalition contributing to preserving humanity's most important information using cryptographic proofs and distributed storage infrastructure. Over 500,000 culturally significant digital artifacts stored on the network.
- Journalism OrganizationscoreThe Internet Archive is responding to publisher blocking by partnering with journalism organizations to train newsrooms in digital preservation, helping journalists maintain access to historical records despite restrictions.
- Vanderbilt University Television News ArchivesupportingStanding on the shoulders of Vanderbilt University's Television News Archive project, the Internet Archive builds upon established TV news archiving methodologies and maintains a collaborative relationship with the advisory board including Vanderbilt representatives.
Scale indicators9 records
Recent moves6 records
Expansion highlights5 records
Internet Archive competitors and assessment
Company assessmentDirect peers
- Project Gutenberg: Volunteer-driven digital library providing free public domain ebooks. Directly comparable to the Internet Archive's Open Library and book digitization efforts, with overlapping collections and shared cultural mission of free knowledge access.
- HathiTrust: Nonprofit digital preservation repository founded by academic and research libraries, preserving 18+ million digitized volumes. Directly comparable to the Internet Archive's book and document preservation mission, with similar copyright challenges and institutional library partnerships.
- Common Crawl: Nonprofit that crawls and archives the open web at massive scale, providing raw datasets to researchers and AI companies. Directly comparable to the Wayback Machine's web archiving mission, though focused on serving datasets rather than public browsing.
- OCLC / WorldCat: Nonprofit library cooperative providing shared cataloging and discovery services. Comparable as nonprofit infrastructure serving libraries worldwide, with overlapping institutional library customer base and metadata management expertise.
Broad incumbents
- Wikimedia Foundation: Nonprofit operating Wikipedia, Wikimedia Commons, and Wikisource. Comparable as a major nonprofit knowledge preservation organization with similar donation-funded model, shared open knowledge mission, and institutional legitimacy, though focused on crowd-sourced content rather than archival preservation.
- Library of Congress: National library of the United States with extensive digital preservation programs. Comparable as a major archive institution with public access mission and digital collections, though funded by federal appropriations rather than donations.
- National Archives and Records Administration (NARA): U.S. federal agency preserving government documents and historical records. Comparable as a major archive institution, particularly relevant given the Internet Archive's Democracy's Library initiative to preserve government data.
- The Library of Congress Digital Collections: Government-operated digital preservation program including historical newspapers, manuscripts, and media. Comparable as a major public digital archive with similar collections spanning books, audio, video, and historical records.
- Google Books: Commercial service from Google that has digitized millions of books. Comparable as a digital book archiving service but operates with commercial intent and publisher partnerships rather than nonprofit preservation mission.
Regional players
- JSTOR: Digital library of academic journals, books, and primary sources. Comparable as a digital preservation and access service for scholarly content, though operating under commercial subscription model with publisher partnerships.
Market position
Strengths4 records
Weaknesses4 records
Competitive moat5 records
Key risks6 records
Key highlights5 records
Customer concentration
Internet Archive social profiles
Digital presenceInternet Archive financial estimates
Financial estimateRevenue estimate
Valuation estimate
Internet Archive leadership team
Management profileNumber of profiles
Profiles6 records
Internet Archive subsidiaries and ownership
Company hierarchySubsidiaries2 records
Internet Archive funding detail
Funding detailFunding overview
Funding rounds1 record
Investors1 record
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Internet Archive M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Internet Archive
What does Internet Archive do?
The Internet Archive is a nonprofit digital library that preserves and provides free public access to a vast collection of digital content including web pages (via the Wayback Machine), books, audio recordings, video, software, and images. It operates one of the world's largest digital preservation infrastructures, storing over 210 petabytes of data and adding approximately 100 terabytes of new material daily. The organization offers free access to most services for individual users, while institutional customers can subscribe to Archive-It for curated web archiving solutions.
Is Internet Archive a public or private company?
Internet Archive is a private company. It is classified as nonprofit foundation owned and is currently operating.
When was Internet Archive founded?
Internet Archive was founded in 1996. It employs 51 to 100 people.
Where is Internet Archive based?
Internet Archive is headquartered in San Francisco, United States, in the North America region.
How does Internet Archive make money?
Four revenue lines are on record. Donations are the primary driver. The others are grants, archive-It Subscription and merchandise.
Who are Internet Archive's main competitors?
Direct peers on record are Project Gutenberg, HathiTrust, Common Crawl and OCLC / WorldCat. Broad incumbents are Wikimedia Foundation, Library of Congress, National Archives and Records Administration (NARA), The Library of Congress Digital Collections and Google Books. JSTOR is listed as a regional player.
Does Internet Archive have an API?
Yes. The Internet Archive provides multiple public APIs including the Wayback Machine API (available at web.archive.org) for looking up archived URLs and retrieving snapshots, and the Save Page Now API for programmatically saving web pages to the archive. Additional developer resources are available at archive.org/developers. The Open Library API provides access to book metadata and lending information. Developer documentation is at archive.org/developers.
What industry is Internet Archive in?
Internet Archive's product category is Digital Library and Web Archiving. Its primary akta.pro industry code is MPAJAMAJ, Digital Archives, Libraries & Institutional Access Platforms, with a secondary code of MPAHAGAF, Cloud Media Archive & Cold Storage Services. Its NAICS code is 51921 and its SIC code is 7370.