DataStax
DataStax provides a real-time data platform for generative AI applications, centered on Astra DB (a serverless Cassandra-based vector database), Langflow (an open-source low-code AI builder), and streaming products. It serves enterprise, digital-native, and platform-engineering teams, and now operates as an IBM company following a February 2025 acquisition.
- Company typePrivate
- Founded2010
- HeadquartersSanta Clara, United States
- Headcount501–1,000
- GTM typeB2B
- OfferingSoftware
What DataStax does
DataStax builds a real-time data platform for generative AI applications, headquartered in Santa Clara, California and founded in 2010. Its product portfolio centers on Astra DB, a serverless NoSQL and vector database built on Apache Cassandra; Hyper-converged Database (HCD), the on-premises counterpart; Astra Streaming, a managed Apache Pulsar service; and the DataStax Enterprise (DSE) data layer. The company acquired Langflow in 2024, an open-source low-code builder for AI applications and multi-agent workflows that has surpassed 100,000 GitHub stars. In October 2024 DataStax launched its AI Platform built with NVIDIA AI, integrating NVIDIA NIM and NeMo Retriever for retrieval-augmented generation. The platform holds HIPAA, SOC 2, ISO 27001, and PCI DSS certifications, and has been recognized as a Leader by both Forrester and Gartner in vector and cloud-native database evaluations.
DataStax's go-to-market combines a freemium developer tier (Astra DB) with enterprise subscriptions for Astra DB, DSE, and HCD, usage-based consumption via cloud marketplaces (AWS, Azure, GCP), and professional services. Pricing is anchored on Astra DB plan tiers, DSE/HCD node-based licensing, and per-capacity-unit (PCU) metering for streaming and cloud consumption. The customer base spans regulated enterprises, digital-native builders, and platform-engineering teams running Gen AI workloads, with Wikimedia Deutschland as a publicly cited production deployment. The company raised approximately $115 million from Goldman Sachs in 2022 at a $1.6 billion valuation and was acquired by IBM in February 2025 for roughly $2.1 billion, after which it operates as "DataStax, an IBM company." The employee count is reported in the 501–1,000 range, distributed across the U.S., EMEA, and Asia following the 2022 China expansion.
DataStax firmographics
Firmographics- Name
- DataStax
- Legal name
- DataStax Inc.
- Website
- https://datastax.com
- Company type
- Private
- Founded year
- 2010
- Operating status
- Acquired
- Headcount range
- 501–1,000 employees
- Short description
- DataStax provides a real-time data platform for generative AI applications, centered on Astra DB (a serverless Cassandra-based vector database), Langflow (an open-source low-code AI builder), and streaming products. It serves enterprise, digital-native, and platform-engineering teams, and now operates as an IBM company following a February 2025 acquisition.
- Ownership category
- akta.pro rank
DataStax industry classification
Industry- Product category
- Cloud Database (Vector/NoSQL)
- NAICS
- Custom Computer Programming Services (541511), Software Publishers (5132), Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (5182)
- SIC
- Services-Computer Programming, Data Processing, Etc. (7370), Services-Prepackaged Software (7372)
- akta.pro primary industry
- Enterprise AI Data & Knowledge Platforms (Vector Databases, Knowledge Graphs) (HDAEANAH)
- akta.pro secondary industries
- Managed Databases (Relational/NoSQL/In-Memory DBaaS) (HDABAAAE), Database-as-a-Service (DBaaS) Platforms (HDAEAAAN), Database Tools & Ecosystem (Replication, Backup/Recovery, HA/DR, Monitoring) (HDAEAAAO), Data Sharing, Data Exchange & Data Marketplace Platforms (HDAEABAH), Data Platform (Unified Data & Analytics) Suites (HDAEABAD)
Keywords
Where DataStax is headquartered
LocationHeadquarters
- HQ city
- Santa Clara
- HQ country
- United States
- HQ region
- North America
Offices6 records
Markets served
DataStax business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Personnel, Technology or R&D, Infrastructure, Marketing or Sales, Operations
Revenue model
- Astra DB Subscription Plans: Recurring subscription revenue from the Astra DB Serverless product with multiple tiers including Free, Standard (Pay-As-You-Go and Subscription), and Enterprise plans. The Marketplace plan (now Standard) is billed through cloud provider marketplaces.
- Provisioned Capacity Units (PCUs): Usage-based revenue from Provisioned Capacity Units (PCUs) with Reserved Capacity Units (RCUs) for committed workloads and Hourly Capacity Units (HCUs) for flexible/auto-scaled capacity. PCUs support complex, latency-sensitive enterprise workloads with dedicated tenant environments.
- DataStax Enterprise (DSE) & Hyper-converged Database (HCD) Licenses: On-premises and private cloud database software licensing for DSE and HCD, including support subscriptions, with DataStax Enterprise Premium edition delivered as DataStax with IBM watsonx.data Premium edition.
- Enterprise Support Services & Professional Services: Premium support plans, consulting, and integration services bundled with enterprise contracts, including contact support, knowledge base access, and DataStax documentation services.
- Cloud Marketplace Consumption: Consumption-based revenue transacted through AWS, Google Cloud, and Microsoft Azure marketplaces where customers pay via the relevant cloud provider for their Astra DB usage.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Freemium | Monthly | Free plan with monthly credits |
| Usage-based | Pay-as-you-go | Standard plan (formerly Marketplace) |
| Subscription | Annual | Enterprise plan |
| Hybrid | Pay-as-you-go | Provisioned Capacity Units (PCUs) |
| Freemium | Monthly | Free trial for Astra DB |
Go-to-market motion5 records
Distribution channels5 records
Marketing channels10 records
DataStax product offering
Product offeringCore offering
DataStax provides a serverless NoSQL and vector database (Astra DB) built on Apache Cassandra, with an on-premises counterpart (Hyper-converged Database, HCD) and an open-source low-code AI builder (Langflow) for prototyping, building, and deploying retrieval-augmented generation and multi-agent AI applications. The DataStax AI Platform, built with NVIDIA AI, integrates these databases with NVIDIA NeMo Retriever and NIM microservices to support enterprise generative AI workloads across on-prem, hybrid, and multi-cloud environments.
Product overview
DataStax, an IBM company, offers a unified AI-ready data platform for managing real-time, unstructured, and multimodal data at scale. The core of the platform is Astra DB (serverless NoSQL vector database built on Apache Cassandra) and HCD (Hyper-converged Database, the on-premises/private cloud counterpart), which together provide vector search, multi-model support (tabular, search, graph), and elastic scalability with near-zero latency. Complementing these are Langflow (an open-source low-code IDE for prototyping, building, and deploying RAG and multi-agent AI applications, with 100,000+ GitHub stars) and the DataStax AI Platform (built with NVIDIA AI, including NeMo Retriever and NIM microservices). Supporting products include Astra Streaming (managed Apache Pulsar event streaming), Mission Control (Kubernetes-based cluster management), Astra CLI (resource management CLI), DataStax Enterprise (DSE) on-premises database platform, OpsCenter (monitoring), DataStax Bulk Loader (DSBulk), DataStax Studio (interactive IDE), Stargate (open-source API gateway), and K8ssandra (Kubernetes operator). As part of IBM watsonx, DataStax delivers vector search, retrieval-augmented generation, and unstructured data management for enterprise AI applications running on-prem, hybrid, or multi-cloud.
Differentiator
Problem solved
Functional benefit
Brands
- Astra DB: DataStax's serverless, multi-cloud NoSQL database built on Apache Cassandra, with native vector search capabilities and integrations for generative AI workloads.
- Hyper-converged Database (HCD)
- Langflow
- Astra Streaming
- DataStax Enterprise (DSE)
- Mission Control
Products and services
- Astra DB Astra DB is a serverless, cloud-native NoSQL database built on Apache Cassandra, providing vector search, multi-model support (tabular, search, graph), and elastic scalability for AI workloads with near-zero latency. Named a Forrester Leader for delivering NoSQL vector search capabilities on cloud.
- Hyper-converged Database (HCD) HCD is the on-premises or private cloud counterpart to Astra DB, delivering the same vector-enabled NoSQL database capabilities for organizations that require on-premise deployment. Delivered as DataStax with IBM watsonx.data Premium edition; available in HCD 2.0, 1.2, and 1.1 versions.
- Langflow Langflow is an open-source low-code visual development environment (100,000+ GitHub stars) for prototyping, building, and deploying retrieval-augmented generation and multi-agent AI applications. Built in Python, designed to work across models, APIs, and databases, and integrating with IBM watsonx Orchestrate as middleware.
- Astra Streaming Astra Streaming is a fully managed event streaming service based on Apache Pulsar that enables efficient data streaming, real-time processing, and CDC for Astra DB Serverless.
- Astra CLI Astra CLI is a command-line interface (v1.0.0 released November 2025) for managing Astra resources — databases, streaming tenants, keyspaces, tokens, organizations — through scripts or the local terminal, with colorized output, extended auto-completions, and Windows support.
- Mission Control Mission Control is a Kubernetes-based platform for managing and orchestrating DSE and Cassandra clusters with declarative resources (MissionControlCluster, CassandraDatacenter), lifecycle manager (LCM) integration, and bundled observability components.
- DataStax AI Platform The DataStax AI Platform, built with NVIDIA AI (including NVIDIA NeMo Retriever and NIM microservices), is a platform-as-a-service for enterprise generative AI development that integrates Astra DB vector search with NVIDIA foundation model tooling.
- DataStax Enterprise (DSE) DSE is the on-premises enterprise database platform based on Apache Cassandra with Advanced Workloads including DSE Search (Apache Solr), DSE Analytics (Apache Spark), and DSE Graph.
- OpsCenter OpsCenter (version 6.8) is a browser-based operations and monitoring tool for DataStax Enterprise (DSE) clusters providing backup/restore, repair, configuration management, and cluster visualization.
- DataStax Bulk Loader (DSBulk) DSBulk is a high-performance bulk loading and unloading utility for Cassandra-based databases supporting CSV, JSON, and other formats for large-scale data migrations and ETL.
- DataStax Studio DataStax Studio is an interactive developer notebook IDE for working with CQL, Gremlin, and Cassandra-based databases, supporting query prototyping, visualization, and code execution.
- Stargate Stargate is an open-source data API gateway for Cassandra that provides Document, REST, GraphQL, and gRPC APIs on top of Cassandra databases.
- K8ssandra K8ssandra is an open-source Kubernetes operator and distribution for Apache Cassandra that provides automated operations, scaling, backup, and observability on Kubernetes clusters.
- CDC for Apache Cassandra CDC for Apache Cassandra provides change data capture functionality for Cassandra databases, capturing row-level changes in real time for streaming into downstream systems.
Quantifiable outcome
- 30-fold increase in query speed and 90% reduction in development time for Wikimedia Deutschland's Wikidata multilingual knowledge graph
- +3 more outcomes
Companies that use DataStax
Customer profileNamed customers1 record
Segments5 records
Ideal customer profiles5 records
DataStax technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration58 records
AI capability10 records
Feature12 records
DataStax partnerships and signals
Strategic signalPartnerships
16 partnerships are on record, tiered flagship and core.
- IBM (International Business Machines)flagshipIBM acquired DataStax in February 2025 for approximately $2.1 billion to complement the watsonx portfolio. DataStax is now an IBM company, bringing Astra DB, HCD, and Langflow to watsonx. The acquisition strengthens IBM's enterprise AI data infrastructure capabilities for unstructured and multimodal data management.
- Wikimedia DeutschlandcoreWikimedia Deutschland launched an AI Knowledge Project in collaboration with DataStax, built with NVIDIA AI. The project integrated DataStax Astra DB on IBM watsonx.data to enhance access to Wikidata's multilingual knowledge graph, achieving a 30-fold increase in query speed and a 90% reduction in development time.
- NVIDIAflagshipDataStax and NVIDIA jointly built the DataStax AI Platform using NVIDIA AI technology. The platform integrates NVIDIA NIM microservices and NeMo Retriever for retrieval-augmented generation. Co-marketed launches have been conducted, including a project with Wikimedia Deutschland using DataStax Astra DB built with NVIDIA AI.
- Jina AIcoreJina AI collaborated with DataStax and Wikimedia Deutschland to launch a semantic search solution for non-profit AI developers. Jina AI is also listed as an integrated embedding provider within Astra DB Serverless (Astra vectorize).
- OpenAIcoreOpenAI is integrated as a vectorize embedding provider within Astra DB Serverless, enabling automatic generation of embeddings for vector operations via Astra vectorize.
- Hugging FacecoreHugging Face is integrated as an external embedding provider within Astra DB Serverless via Astra vectorize, available in both Dedicated and Serverless integration modes.
- Microsoft Azure OpenAIcoreAzure OpenAI is integrated as a vectorize embedding provider within Astra DB Serverless for automated embedding generation.
- LangChaincoreLangChain is integrated with Astra DB Serverless and is the foundational framework for Langflow (DataStax's low-code AI builder). Available in Python and JavaScript variants.
- Amazon Web Services (AWS)coreAstra DB Serverless is available on AWS with integrations to AWS Bedrock, AWS SageMaker, AWS Glue, AWS Lambda, and AWS PrivateLink. Billing is supported via AWS Marketplace.
- Google Cloud PlatformcoreAstra DB Serverless is available on Google Cloud with integrations to Google Cloud Functions, Google Dataflow, Google Vertex AI, and Google Cloud Private Service Connect. Billing is supported via Google Cloud Marketplace.
- Microsoft AzurecoreAstra DB Serverless is available on Microsoft Azure with integrations to Azure Functions and Azure Private Link. Billing is supported via Azure Marketplace.
- Amazon BedrockcoreAmazon Bedrock is integrated with Astra DB Serverless for managed foundation model access.
- Google Vertex AIcoreGoogle Vertex AI is integrated with Astra DB Serverless as an Extension and for Search and chat capabilities.
- GleancoreDataStax delivered Glean and Unstructured integrations to the AI Platform at the RAG++ Event in NYC, enabling enterprise search and RAG capabilities.
- Unstructured.iocoreUnstructured Serverless integration is available with Astra DB Serverless for unstructured data ingestion, delivered at the RAG++ Event in NYC.
- Model Context Protocol (MCP)coreAstra DB MCP server is available as a replacement for the deprecated Astra DB GitHub Copilot extension, enabling AI agent integrations with Astra DB.
Scale indicators14 records
Recent moves7 records
Expansion highlights6 records
DataStax competitors and assessment
Company assessmentDirect peers
- Pinecone: Pinecone is a managed vector database purpose-built for production RAG and semantic search workloads. It competes head-to-head with Astra DB in the vector database category and is named alongside DataStax in market commentary on the segment.
- Weaviate: Weaviate is an open-source vector database with hybrid search and modular AI-native capabilities. It is a directly comparable peer to Astra DB on vector search, RAG, and AI application workloads.
- Qdrant: Qdrant is an open-source vector similarity search engine written in Rust, focused on high-performance vector and hybrid retrieval for AI applications. It competes directly with Astra DB's vector search capabilities.
- Zilliz (Milvus): Zilliz is the commercial entity behind the open-source Milvus vector database. Both Milvus and Zilliz Cloud are explicitly named alongside DataStax in the vector database market commentary, making it a direct peer.
- Chroma: Chroma is an open-source embedding database widely used for LLM application development and prototyping. It overlaps with Astra DB on vector storage for AI applications and competes for the same AI/ML developer mindshare as Langflow-backed Astra workflows.
Broad incumbents
- MongoDB: MongoDB is a leading general-purpose NoSQL document database that has added native vector search capabilities via Atlas Vector Search. It is a broader incumbent competing with Astra DB for enterprise document and AI workloads across hybrid cloud.
- Couchbase: Couchbase is a distributed NoSQL document and key-value database with vector search capabilities targeting enterprise and edge workloads. Its distributed architecture and Cassandra-like operational profile make it a comparable broad incumbent to DataStax in enterprise NoSQL.
- Elastic: Elastic is a search and analytics platform with dense vector retrieval, hybrid search, and RAG capabilities integrated with Elasticsearch. It is a broader incumbent competing for enterprise search and AI infrastructure budgets alongside Astra DB.
- Redis: Redis is an in-memory data store that has added vector search (RedisVL) and is positioned for low-latency AI application use cases. Its presence in real-time application stacks makes it a broad incumbent peer for DataStax's low-latency AI data infrastructure.
Emerging players
- Confluent: Confluent is the commercial provider of Apache Kafka and a streaming data platform, comparable to DataStax's Astra Streaming built on Apache Pulsar. It competes in the same managed event-streaming category, including for CDC and AI data pipelines.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat6 records
Key risks6 records
Key highlights7 records
Customer concentration
DataStax social profiles
Digital presenceDataStax compliance and trust
Trust signalCompliance4 records
DataStax financial estimates
Financial estimateRevenue estimate
Valuation estimate
DataStax leadership team
Management profileNumber of profiles
Profiles14 records
DataStax subsidiaries and ownership
Company hierarchySubsidiaries1 record
DataStax funding detail
Funding detailFunding overview
Funding rounds10 records
Investors23 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
DataStax M&A and investment
M&A and investmentM&A6 records
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about DataStax
What does DataStax do?
DataStax provides a serverless NoSQL and vector database (Astra DB) built on Apache Cassandra, with an on-premises counterpart (Hyper-converged Database, HCD) and an open-source low-code AI builder (Langflow) for prototyping, building, and deploying retrieval-augmented generation and multi-agent AI applications. The DataStax AI Platform, built with NVIDIA AI, integrates these databases with NVIDIA NeMo Retriever and NIM microservices to support enterprise generative AI workloads across on-prem, hybrid, and multi-cloud environments.
Is DataStax a public or private company?
DataStax is a private company. It is classified as corporate owned and is currently acquired.
When was DataStax founded?
DataStax was founded in 2010. It employs 501 to 1,000 people.
Where is DataStax based?
DataStax is headquartered in Santa Clara, United States, in the North America region.
How does DataStax make money?
Five revenue lines are on record. Astra DB Subscription Plans are the primary driver. The others are provisioned Capacity Units (PCUs), dataStax Enterprise (DSE) & Hyper-converged Database (HCD) Licenses, enterprise Support Services & Professional Services and cloud Marketplace Consumption.
Who are DataStax's main competitors?
Direct peers on record are Pinecone, Weaviate, Qdrant, Zilliz (Milvus) and Chroma. Broad incumbents are MongoDB, Couchbase, Elastic and Redis. Confluent is listed as an emerging player.
Does DataStax have an API?
Yes. DataStax offers multiple APIs including the Data API (schema-less, document-based, JSON API for Astra DB Serverless vector databases with Python, TypeScript, Java, and C# client libraries) and the DevOps API (administrative REST API for managing Astra DB resources, organizations, databases, keyspaces, access lists, and metrics). The Data API enables developers to build production generative AI and retrieval-augmented generation (RAG) applications by inserting, finding, updating, and deleting documents in collections and rows in tables. The DevOps API v2 provides programmatic access to provision and manage Astra DB Serverless databases. Authentication is handled via application tokens or SSO. Clients support table-based structured data (in preview) and hybrid search. Developer documentation is at docs.datastax.com/en/astra-db-serverless/api-reference/dataapiclient.html.
What industry is DataStax in?
DataStax's product category is Cloud Database (Vector/NoSQL). Its primary akta.pro industry code is HDAEANAH, Enterprise AI Data & Knowledge Platforms (Vector Databases, Knowledge Graphs), with a secondary code of HDABAAAE, Managed Databases (Relational/NoSQL/In-Memory DBaaS). Its NAICS code is 541511 and its SIC code is 7370.