Hadoop
Apache Hadoop is an open-source distributed-computing framework under the Apache Software Foundation that provides HDFS storage, YARN resource management, and MapReduce processing across clusters, with cloud connectors and an ecosystem spanning Spark, Hive, HBase, Ozone, and Submarine for global enterprise and research use under Apache License 2.0.
- Company typePrivate
- Founded2006
- HeadquartersBaltimore, United States
- Headcount11–50
- GTM typeB2B
- OfferingSoftware
What Hadoop does
Apache Hadoop is an open-source software framework originally established in 2006 inside the Apache Software Foundation and maintained as a community project under the Apache License 2.0. It provides four core modules — Hadoop Common, the Hadoop Distributed File System (HDFS), Hadoop YARN, and Hadoop MapReduce — that together enable reliable, scalable distributed processing of large data sets across clusters of machines, with fault tolerance handled at the application layer rather than via hardware redundancy.
The framework is complemented by a substantial ecosystem of related Apache projects including HBase (distributed NoSQL), Hive (data warehouse), Spark (general compute engine), Pig (data-flow language), Ozone (object store), Ambari (cluster management), ZooKeeper (coordination), Avro (serialization), Cassandra, Tez (DAG execution), Mahout (machine learning), and Submarine (unified AI/ML platform). Native connectors exist for major cloud object stores (AWS S3, Azure ABFS/ADLS, Aliyun OSS, Tencent COS, Huawei OBS), allowing the same HDFS/YARN abstractions to operate against cloud-resident data. The framework scales from a single server to thousands of machines and offers a Java API plus a WebHDFS REST API.
The business model is purely open-source distribution: the project generates no direct revenue, has no commercial license, no pricing tiers, and no enterprise sales motion; it is distributed via Apache download mirrors, Maven Central, and the Apache GitBox repository. All investment in development is contributed by volunteer committers, PMC members, and sponsoring organizations, with the foundation itself operating as a U.S.-incorporated nonprofit. The user base spans enterprise data teams, data engineering groups, and academic/research institutions globally.
Hadoop firmographics
Firmographics- Name
- Hadoop
- Legal name
- Apache Hadoop
- Website
- https://hadoop.apache.org
- Company type
- Private
- Founded year
- 2006
- Operating status
- Operating
- Headcount range
- 11–50 employees
- Short description
- Apache Hadoop is an open-source distributed-computing framework under the Apache Software Foundation that provides HDFS storage, YARN resource management, and MapReduce processing across clusters, with cloud connectors and an ecosystem spanning Spark, Hive, HBase, Ozone, and Submarine for global enterprise and research use under Apache License 2.0.
- Ownership category
- akta.pro rank
Hadoop industry classification
Industry- Product category
- Distributed Computing Framework
- NAICS
- Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (5182)
- SIC
- Services-Prepackaged Software (7372)
- akta.pro primary industry
- NoSQL Database Management Systems (HDAEAAAB)
- akta.pro secondary industries
- Relational Database Management Systems (RDBMS) (HDAEAAAA), Cloud Networking Core Services (VPC/VNet, Load Balancing, DNS, NAT) (HDABAAAJ)
Keywords
Where Hadoop is headquartered
LocationHeadquarters
- HQ city
- Baltimore
- HQ country
- United States
- HQ region
- North America
Markets served
Hadoop business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Technology or R&D, Personnel, Operations, Infrastructure
Distribution channels3 records
Marketing channels4 records
Hadoop product offering
Product offeringCore offering
Apache Hadoop is an open-source software framework that enables the distributed processing and storage of large data sets across clusters of computers using simple programming models. It scales from a single server to thousands of machines, delivers high availability through application-layer fault tolerance rather than hardware redundancy, and is composed of four core modules: Hadoop Common, HDFS, YARN, and MapReduce. The framework is released under the Apache License 2.0 and is freely downloadable for research and production use.
Product overview
Apache Hadoop is an open-source software framework consisting of four core modules (Hadoop Common, HDFS, YARN, and MapReduce) that work together to provide reliable, scalable, distributed computing. The framework is designed to scale from single servers to thousands of machines, each offering local computation and storage, with fault tolerance handled at the application layer rather than relying on hardware. The Hadoop ecosystem extends the core framework with numerous related projects including HBase (NoSQL database), Hive (data warehouse with SQL querying), Spark (fast compute engine), Ozone (object store), Pig (data-flow language), Ambari (cluster management), ZooKeeper (coordination service), Mahout (machine learning), Submarine (unified AI platform for ML/DL workloads), and Tez (generalized data-flow framework). These components form a comprehensive big data processing platform used for research and production across a wide variety of organizations.
Differentiator
Problem solved
Functional benefit
Brands
- Hadoop Common: The common utilities that support the other Hadoop modules.
- Hadoop Distributed File System (HDFS)
- Hadoop YARN
- Hadoop MapReduce
Products and services
- Apache Hadoop Open-source software framework for reliable, scalable, distributed computing. Enables distributed processing of large data sets across clusters using simple programming models, scales from single servers to thousands of machines, and provides fault tolerance at the application layer for enterprise data teams, research institutions, and data engineering teams.
Quantifiable outcome
- Scales to thousands of machines with local computation and storage per node
Companies that use Hadoop
Customer profileSegments3 records
Ideal customer profiles3 records
Hadoop technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration10 records
Feature5 records
Hadoop partnerships and signals
Strategic signalScale indicators3 records
Recent moves6 records
Expansion highlights4 records
Hadoop competitors and assessment
Company assessmentBroad incumbents
- Databricks: Unified data analytics platform built on Apache Spark, often deployed alongside or as an alternative to Hadoop for data engineering, ML, and analytics workloads.
- Amazon EMR: Managed Hadoop framework service on AWS that monetizes Hadoop as cloud infrastructure, supporting Hadoop, Spark, HBase, and Presto on EC2 clusters.
- Cloudera: Commercial Hadoop distribution provider (formed from Cloudera-Hortonworks merger) offering enterprise support, management tools, and data platform services built on Apache Hadoop.
- Snowflake: Cloud-native data warehouse platform that competes with Hive-on-Hadoop and other Hadoop-based SQL query workloads, offering simpler managed experience and separation of storage and compute.
- Google Cloud Dataproc: Managed Spark and Hadoop service on Google Cloud, providing similar managed Hadoop capabilities as AWS EMR with integration to GCP storage and BigQuery.
Direct peers
- Apache Flink: Distributed stream and batch processing framework that competes directly with Hadoop MapReduce/YARN for large-scale data processing workloads, particularly in real-time streaming use cases.
- Apache Kafka: Distributed streaming platform commonly paired with Hadoop in modern data architectures, providing real-time event ingestion that complements Hadoop's batch processing strengths.
- Apache Spark: Fast general compute engine that runs on Hadoop YARN and provides ETL, machine learning, stream processing, and graph computation. Same Apache ecosystem, often deployed together on shared Hadoop infrastructure.
Emerging players
- Dremio: Data lakehouse platform that provides SQL querying and semantic layer on data lake storage, competing with Hive and other Hadoop-based query engines for modern lakehouse workloads.
- Trino (formerly Presto): Distributed SQL query engine for big data that competes with Hive for interactive querying on Hadoop data lakes, with growing adoption for federated queries across data sources.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat5 records
Key risks5 records
Key highlights7 records
Customer concentration
Hadoop social profiles
Digital presenceHadoop financial estimates
Financial estimateRevenue estimate
Valuation estimate
Hadoop leadership team
Management profileNumber of profiles
Hadoop funding detail
Funding detailFunding overview
Funding rounds
Investors
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Hadoop M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Hadoop
What does Hadoop do?
Apache Hadoop is an open-source software framework that enables the distributed processing and storage of large data sets across clusters of computers using simple programming models. It scales from a single server to thousands of machines, delivers high availability through application-layer fault tolerance rather than hardware redundancy, and is composed of four core modules: Hadoop Common, HDFS, YARN, and MapReduce. The framework is released under the Apache License 2.0 and is freely downloadable for research and production use.
Is Hadoop a public or private company?
Hadoop is a private company. It is classified as nonprofit foundation owned and is currently operating.
When was Hadoop founded?
Hadoop was founded in 2006. It employs 11 to 50 people.
Where is Hadoop based?
Hadoop is headquartered in Baltimore, United States, in the North America region.
Who are Hadoop's main competitors?
Broad incumbents on record are Databricks, Amazon EMR, Cloudera, Snowflake and Google Cloud Dataproc. Direct peers are Apache Flink, Apache Kafka and Apache Spark. Emerging players are Dremio and Trino (formerly Presto).
Does Hadoop have an API?
Yes. Hadoop provides a Java API for programmatic access and WebHDFS (REST API) for HTTP-based access to HDFS. The Java API docs are available at https://hadoop.apache.org/docs/r3.4.0/api/. WebHDFS provides REST endpoints for file operations including read, write, create, delete, and other filesystem operations. The API allows developers to build applications that interact with Hadoop's distributed file system and YARN resource management. Developer documentation is at hadoop.apache.org/docs/r3.4.0/api.
What industry is Hadoop in?
Hadoop's product category is Distributed Computing Framework. Its primary akta.pro industry code is HDAEAAAB, NoSQL Database Management Systems, with a secondary code of HDAEAAAA, Relational Database Management Systems (RDBMS). Its NAICS code is 5182 and its SIC code is 7372.