Apache Gravitino
Apache Gravitino is an Apache Software Foundation Top-Level Project providing a federated, geo-distributed metadata lake for unified management of data and AI assets across 15+ heterogeneous sources, with native connectors for Trino, Spark, Flink, Iceberg, and vector data.
- Company typePrivate
- Founded2024
- Headquarters—
- Headcount—
- GTM typeB2B
- OfferingSoftware
What Apache Gravitino does
Apache Gravitino is an Apache Software Foundation Top-Level Project providing a high-performance, geo-distributed, federated metadata lake for unified management of data and AI assets across heterogeneous sources. The platform, written in Java and requiring JDK 17, exposes REST APIs and Java/Python SDKs, with connectors spanning 15+ systems: relational databases (MySQL, PostgreSQL, Doris, StarRocks, ClickHouse, Hologres, OceanBase), lakehouse formats (Apache Iceberg, Apache Hudi, Apache Paimon, Delta Lake), messaging (Apache Kafka), filesets (HDFS, S3, GCS, OSS), and AI/ML assets (models, features, vector data). Its distinguishing technical architecture is direct, bidirectional metadata management—changes in Gravitino propagate to underlying systems and vice versa—contrasted with traditional passive collection. Additional capabilities include a Model Context Protocol (MCP) server for AI agent integration, a Table Maintenance Service for automated lakehouse optimization, credential vending for secure multi-cloud access, Apache Ranger integration for unified RBAC, and standalone Iceberg REST and Lance REST catalog services.
The project entered the Apache Incubator in June 2024 and graduated to Top-Level Project status on June 3, 2025. Since then it has shipped eight releases through version 1.3.0 (June 2026), with named production adopters including Uber (multi-cloud AI clusters) and Pinterest (Iceberg REST Catalog in production); Datastrato has publicly committed support. As a community-led open-source project distributed under Apache License 2.0, Gravitino has no direct revenue model—commercial adoption occurs through self-hosted enterprise deployments and third-party support providers. The community has grown to 2,600+ GitHub stars (130%+ YoY) and 40+ unique developers per major release, with visitors from 103 countries reaching documentation in the second half of 2024.
Apache Gravitino is governed as an open-source community project under the Apache Software Foundation, a US 501(c)(3) nonprofit, with no parent-company control beyond ASF oversight. Go-to-market is community-led, relying on GitHub, ASF Slack (#gravitino), developer mailing lists, technical documentation, blog posts, and conference appearances at Community Over Code (NA and Asia) and QCon Shanghai. Distribution is frictionless via Apache download mirrors, Docker Hub, GitHub, and versioned Trino connector packages.
Apache Gravitino firmographics
Firmographics- Name
- Apache Gravitino
- Legal name
- Apache Gravitino
- Website
- https://gravitino.apache.org
- Company type
- Private
- Founded year
- 2024
- Operating status
- Operating
- Short description
- Apache Gravitino is an Apache Software Foundation Top-Level Project providing a federated, geo-distributed metadata lake for unified management of data and AI assets across 15+ heterogeneous sources, with native connectors for Trino, Spark, Flink, Iceberg, and vector data.
- Ownership category
- akta.pro rank
Apache Gravitino industry classification
Industry- Product category
- Metadata Management
- NAICS
- Computer Systems Design and Related Services (54151), Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (51821), Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (518)
- SIC
- Services-Prepackaged Software (7372), Services-Computer Programming, Data Processing, Etc. (7370), Services-Computer Integrated Systems Design (7373)
- akta.pro primary industry
- Configuration Management Database (CMDB) Platforms (HDAEALAE)
- akta.pro secondary industries
- Governance, Risk & Compliance (GRC) Platforms (BPAEAPAA), Cloud Compliance, Audit & Continuous Controls Monitoring (CCM/GRC) (HDABAHAI)
Keywords
Apache Gravitino business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Technology or R&D, Personnel, Infrastructure, Operations
Revenue model
- Open Source Distribution: Apache Gravitino is an open-source project distributed under the Apache License 2.0. There is no direct revenue model; the project is maintained by community contributors and sponsored by the Apache Software Foundation. Commercial adoption occurs through organizations using the open-source software independently or through commercial support providers in the ecosystem.
Go-to-market motion1 record
Distribution channels4 records
Marketing channels6 records
Apache Gravitino product offering
Product offeringCore offering
Apache Gravitino is a high-performance, geo-distributed, federated metadata lake that provides unified metadata management across heterogeneous data and AI assets — including relational databases, file/HDFS stores, streaming systems, and AI models — accessible from multiple query/compute engines such as Trino, Spark, Flink, and Daft. It also offers standalone REST catalog services (Apache Iceberg REST, Lance REST), a Model Context Protocol (MCP) server, and a Table Maintenance Service for unified governance, access control, and asset management.
Product overview
Apache Gravitino is a unified, federated metadata lake platform for managing data and AI assets across heterogeneous sources. The core Gravitino server (with REST API and Java/Python SDKs) provides unified metadata management, governance, and access control for relational databases, file stores, and event streams. Key products include specialized REST catalog services for Apache Iceberg and Lance (vector data), connectors for Trino/Spark/Flink/Daft query engines, Web UI and CLI management interfaces, and an MCP server for AI agent integration. The platform supports multi-engine access (Spark, Trino, Flink), geo-distributed deployment, credential vending for secure cloud storage access, and metadata-driven automation via statistics, policies, and job systems.
Differentiator
Problem solved
Functional benefit
Products and services
- Apache Gravitino (Core Server) High-performance, geo-distributed, federated metadata lake and unified catalog server that manages metadata for data and AI assets across relational databases, file stores, and streaming systems, with multi-engine access for Trino, Spark, Flink, and Daft.
- Gravitino Iceberg REST Catalog Service Standards-based, Apache Iceberg REST Catalog-compatible service for managing Iceberg tables across engines such as Trino and Spark.
- Gravitino Lance REST Service REST service for cataloging and managing Apache Lance datasets, exposing Lance-format tables to compatible engines.
- Gravitino MCP Server Model Context Protocol server that exposes Gravitino catalog, governance, and AI-asset metadata to LLM-based agents and external tools.
- Gravitino Table Maintenance Service (TMS / Optimizer) Standalone service that runs optimization jobs such as compaction and cleanup on tables registered in Gravitino.
Quantifiable outcome
- 130%+ GitHub star growth year-over-year
- +2 more outcomes
Companies that use Apache Gravitino
Customer profileNamed customers3 records
Segments3 records
Ideal customer profiles3 records
Apache Gravitino technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration26 records
AI capability4 records
Feature15 records
Apache Gravitino partnerships and signals
Strategic signalPartnerships
One partnership is on record.
- Apache Software FoundationflagshipApache Gravitino is an Apache Top-Level Project (TLP) under the Apache Software Foundation. ASF provides governance, infrastructure, licensing (Apache License 2.0), and vendor-neutral oversight. The project entered Apache Incubator in June 2024 and graduated to TLP status on June 3, 2025. ASF provides the Foundation, License, Events, Privacy Policy, Security, Sponsorship, and Thanks infrastructure.
Scale indicators5 records
Recent moves1 record
Expansion highlights7 records
Apache Gravitino competitors and assessment
Company assessmentBroad incumbents
- Collibra: Enterprise data intelligence platform with strong governance, lineage, and data quality capabilities. Adjacent competitor to Gravitino for unified data governance workloads in large regulated enterprises.
- Alation: Established enterprise data catalog and governance vendor with deep Fortune 500 penetration. Competes with Gravitino on data discovery and governance, though Gravitino's edge is in lakehouse federation and AI asset management rather than Alation's data stewardship focus.
- AWS Glue Data Catalog: Hyperscaler-native metadata catalog integrated with the broader AWS analytics stack (Athena, EMR, Redshift, Lake Formation). Competes with Gravitino on federated metadata access, especially for AWS-centric enterprises, but lacks Gravitino's multi-engine and multi-cloud federation.
Direct peers
- OpenMetadata: Open-source metadata and governance platform covering data discovery, lineage, and collaboration. Competes with Gravitino for unified metadata management workloads in enterprise data platforms with a similar community-driven model.
- Apache Polaris: Apache Incubator project from Snowflake providing an Iceberg REST catalog. The closest sibling project under ASF, competing with Gravitino's Iceberg REST catalog service for the same open-table-format catalog workload.
- DataHub: Open-source metadata platform originally built at LinkedIn. Directly comparable to Gravitino as a unified data catalog with discovery, lineage, and governance; both pursue open-source, community-led distribution models targeting enterprise data platforms.
- Unity Catalog: Open data catalog from Databricks with native support for Iceberg, Delta, and ML assets. Most direct competitor to Gravitino's federated Iceberg REST catalog and AI asset management capabilities, though backed by Databricks rather than vendor-neutral governance.
Emerging players
- Atlan: Active metadata platform with strong focus on data discovery, governance, and collaboration built atop open standards. Competes for the same enterprise data catalog workload with a more productized, commercial-led approach.
- LakeFS: Open-source data lake version control and management platform. Adjacent to Gravitino's metadata and fileset management, with overlapping concern for managing data across S3/GCS/Azure storage layers in lakehouse architectures.
- Project Nessie: Open-source transactional catalog and data lake version control from the Iceberg community (originally Cloudera). Comparable in scope to Gravitino's Iceberg REST catalog service for federated lakehouse metadata access.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat6 records
Key risks6 records
Key highlights7 records
Customer concentration
Apache Gravitino social profiles
Digital presenceApache Gravitino compliance and trust
Trust signalCompliance1 record
Apache Gravitino financial estimates
Financial estimateRevenue estimate
Valuation estimate
Apache Gravitino leadership team
Management profileNumber of profiles
Profiles6 records
Apache Gravitino funding detail
Funding detailFunding overview
Funding rounds
Investors
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Apache Gravitino M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Apache Gravitino
What does Apache Gravitino do?
Apache Gravitino is a high-performance, geo-distributed, federated metadata lake that provides unified metadata management across heterogeneous data and AI assets — including relational databases, file/HDFS stores, streaming systems, and AI models — accessible from multiple query/compute engines such as Trino, Spark, Flink, and Daft. It also offers standalone REST catalog services (Apache Iceberg REST, Lance REST), a Model Context Protocol (MCP) server, and a Table Maintenance Service for unified governance, access control, and asset management.
Is Apache Gravitino a public or private company?
Apache Gravitino is a private company. It is classified as nonprofit foundation owned and is currently operating.
When was Apache Gravitino founded?
Apache Gravitino was founded in 2024.
How does Apache Gravitino make money?
One revenue line is on record: open Source Distribution.
Who are Apache Gravitino's main competitors?
Broad incumbents on record are Collibra, Alation and AWS Glue Data Catalog. Direct peers are OpenMetadata, Apache Polaris, DataHub and Unity Catalog. Emerging players are Atlan, LakeFS and Project Nessie.
Does Apache Gravitino have an API?
Yes. Gravitino provides REST APIs for managing metalakes, relational metadata, fileset metadata, messaging metadata, model metadata, user-defined functions, tags, policies, and jobs. It also exposes Java SDK and Python SDK for client-side access. The API enables developers to programmatically manage metadata across heterogeneous data sources including relational databases, file stores, and event streams. Developer documentation is at gravitino.apache.org/docs/1.3.0/api/rest/gravitino-rest-api.
What industry is Apache Gravitino in?
Apache Gravitino's product category is Metadata Management. Its primary akta.pro industry code is HDAEALAE, Configuration Management Database (CMDB) Platforms, with a secondary code of BPAEAPAA, Governance, Risk & Compliance (GRC) Platforms. Its NAICS code is 54151 and its SIC code is 7372.