Developer docs
API playgroundTry for free, no card

Search company profiles

DataStax

Full company profile

uuid00004nw

Namestring
DataStax
Legal namestring
DataStax Inc.
Websiteurl
datastax.com
Company typeenum
Private
Founded yearint
2010
Descriptiontext

DataStax builds a real-time data platform for generative AI applications, headquartered in Santa Clara, California and founded in 2010. Its product portfolio centers on Astra DB, a serverless NoSQL and vector database built on Apache Cassandra; Hyper-converged Database (HCD), the on-premises counterpart; Astra Streaming, a managed Apache Pulsar service; and the DataStax Enterprise (DSE) data layer. The company acquired Langflow in 2024, an open-source low-code builder for AI applications and multi-agent workflows that has surpassed 100,000 GitHub stars. In October 2024 DataStax launched its AI Platform built with NVIDIA AI, integrating NVIDIA NIM and NeMo Retriever for retrieval-augmented generation. The platform holds HIPAA, SOC 2, ISO 27001, and PCI DSS certifications, and has been recognized as a Leader by both Forrester and Gartner in vector and cloud-native database evaluations.

DataStax's go-to-market combines a freemium developer tier (Astra DB) with enterprise subscriptions for Astra DB, DSE, and HCD, usage-based consumption via cloud marketplaces (AWS, Azure, GCP), and professional services. Pricing is anchored on Astra DB plan tiers, DSE/HCD node-based licensing, and per-capacity-unit (PCU) metering for streaming and cloud consumption. The customer base spans regulated enterprises, digital-native builders, and platform-engineering teams running Gen AI workloads, with Wikimedia Deutschland as a publicly cited production deployment. The company raised approximately $115 million from Goldman Sachs in 2022 at a $1.6 billion valuation and was acquired by IBM in February 2025 for roughly $2.1 billion, after which it operates as "DataStax, an IBM company." The employee count is reported in the 501–1,000 range, distributed across the U.S., EMEA, and Asia following the 2022 China expansion.

Short descriptiontext

DataStax provides a real-time data platform for generative AI applications, centered on Astra DB (a serverless Cassandra-based vector database), Langflow (an open-source low-code AI builder), and streaming products. It serves enterprise, digital-native, and platform-engineering teams, and now operates as an IBM company following a February 2025 acquisition.

Operating statusenum
Acquired
Ownership categoryenum
Headcount rangeband
501–1,000
akta.pro rankint
HeadquartersSanta Clara, United States
HQ citystring
Santa Clara
HQ countrystring
United States
HQ regionstring
North America
Markets served

Serves global market

Offices6 records

Each record includes

City, Country, Type, Description, Source

Keyword5 values
vector database, NoSQL database, retrieval-augmented generation, AI data platform, generative AI infrastructure
Industry6 codes
1Enterprise AI Data & Knowledge Platforms (Vector Databases, Knowledge Graphs)
CodeHDAEANAHPrimaryYes
2Managed Databases (Relational/NoSQL/In-Memory DBaaS)
CodeHDABAAAEPrimaryNo
3Database-as-a-Service (DBaaS) Platforms
CodeHDAEAAANPrimaryNo
4Database Tools & Ecosystem (Replication, Backup/Recovery, HA/DR, Monitoring)
CodeHDAEAAAOPrimaryNo
5Data Sharing, Data Exchange & Data Marketplace Platforms
CodeHDAEABAHPrimaryNo
6Data Platform (Unified Data & Analytics) Suites
CodeHDAEABADPrimaryNo
NAICS code3 codes
  • Custom Computer Programming Services541511
  • Software Publishers5132
  • Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services5182
SIC code2 codes
  • Services-Computer Programming, Data Processing, Etc.7370
  • Services-Prepackaged Software7372
Product category
Cloud Database (Vector/NoSQL)
GTM motion5 records

Each record includes

Type, Description, Source

Revenue model5 records
1Astra DB Subscription Plans
TypeSubscription Recurring
Description

Recurring subscription revenue from the Astra DB Serverless product with multiple tiers including Free, Standard (Pay-As-You-Go and Subscription), and Enterprise plans. The Marketplace plan (now Standard) is billed through cloud provider marketplaces.

docs.datastax.com
2Provisioned Capacity Units (PCUs)
TypeUsage Based
Description

Usage-based revenue from Provisioned Capacity Units (PCUs) with Reserved Capacity Units (RCUs) for committed workloads and Hourly Capacity Units (HCUs) for flexible/auto-scaled capacity. PCUs support complex, latency-sensitive enterprise workloads with dedicated tenant environments.

docs.datastax.com
3DataStax Enterprise (DSE) & Hyper-converged Database (HCD) Licenses
TypeSubscription Recurring
Description

On-premises and private cloud database software licensing for DSE and HCD, including support subscriptions, with DataStax Enterprise Premium edition delivered as DataStax with IBM watsonx.data Premium edition.

ibm.com
4Enterprise Support Services & Professional Services
TypeProfessional Services
Description

Premium support plans, consulting, and integration services bundled with enterprise contracts, including contact support, knowledge base access, and DataStax documentation services.

ibm.com
5Cloud Marketplace Consumption
TypeUsage Based
Description

Consumption-based revenue transacted through AWS, Google Cloud, and Microsoft Azure marketplaces where customers pay via the relevant cloud provider for their Astra DB usage.

docs.datastax.com
Marketing channels10 records

Each record includes

Title, Type, Stage, Description, Source

Distribution channels5 records

Each record includes

Title, Type, Scope, Target buyer, Description, Source

Cost components5 values
Personnel, Technology or R&D, Infrastructure, Marketing or Sales, Operations
Pricing details5 tiers
1Free plan with monthly credits
ModelFreemiumBilling cadenceMonthly
Notes

Free tier provides monthly credits. Databases are suspended if monthly credits run out; can be reactivated by upgrading plan or waiting for credits to refresh.

docs.datastax.com
2Standard plan (formerly Marketplace)
ModelUsage-basedBilling cadencePay-as-you-go
Notes

Includes Pay-As-You-Go and Subscription options. Billed via cloud provider marketplaces (AWS Marketplace, Google Cloud Marketplace, Azure Marketplace).

docs.datastax.com
3Enterprise plan
ModelSubscriptionBilling cadenceAnnual
Notes

Enterprise tier includes enterprise organization management (GA), PCU groups, enterprise tokens, enterprise administration roles, enterprise-wide usage reports, and access to features like Astra DB Sideloader (GA).

docs.datastax.com
4Provisioned Capacity Units (PCUs)
ModelHybridBilling cadencePay-as-you-go
Notes

PCUs are sold as Reserved Capacity Units (RCUs) for committed workloads or Hourly Capacity Units (HCUs) for flexible usage. Maximum 10 PCUs per group; additional capacity can be added via additional PCU groups. PCU pricing depends on subscription options and actual hourly usage.

docs.datastax.com
5Free trial for Astra DB
ModelFreemiumBilling cadenceMonthly
Notes

Free trial available for Astra DB Serverless via signup at astra.datastax.com.

ibm.com
GTM typeB2B
B2B
Offering typeSoftware
Software
Brand1 of 6 records shown
1Astra DB
Description

DataStax's serverless, multi-cloud NoSQL database built on Apache Cassandra, with native vector search capabilities and integrations for generative AI workloads.

ibm.com
+5 more records
Core offering1 text field

DataStax provides a serverless NoSQL and vector database (Astra DB) built on Apache Cassandra, with an on-premises counterpart (Hyper-converged Database, HCD) and an open-source low-code AI builder (Langflow) for prototyping, building, and deploying retrieval-augmented generation and multi-agent AI applications. The DataStax AI Platform, built with NVIDIA AI, integrates these databases with NVIDIA NeMo Retriever and NIM microservices to support enterprise generative AI workloads across on-prem, hybrid, and multi-cloud environments.

Differentiator
Functional benefit
Problem solved
Quantifiable outcome1 of 4 values shown
  • 30-fold increase in query speed and 90% reduction in development time for Wikimedia Deutschland's Wikidata multilingual knowledge graph
+3 more records
Product overview1 text field

DataStax, an IBM company, offers a unified AI-ready data platform for managing real-time, unstructured, and multimodal data at scale. The core of the platform is Astra DB (serverless NoSQL vector database built on Apache Cassandra) and HCD (Hyper-converged Database, the on-premises/private cloud counterpart), which together provide vector search, multi-model support (tabular, search, graph), and elastic scalability with near-zero latency. Complementing these are Langflow (an open-source low-code IDE for prototyping, building, and deploying RAG and multi-agent AI applications, with 100,000+ GitHub stars) and the DataStax AI Platform (built with NVIDIA AI, including NeMo Retriever and NIM microservices). Supporting products include Astra Streaming (managed Apache Pulsar event streaming), Mission Control (Kubernetes-based cluster management), Astra CLI (resource management CLI), DataStax Enterprise (DSE) on-premises database platform, OpsCenter (monitoring), DataStax Bulk Loader (DSBulk), DataStax Studio (interactive IDE), Stargate (open-source API gateway), and K8ssandra (Kubernetes operator). As part of IBM watsonx, DataStax delivers vector search, retrieval-augmented generation, and unstructured data management for enterprise AI applications running on-prem, hybrid, or multi-cloud.

Product and service14 records
1Astra DB
CategoryCore product (serverless NoSQL vector database)
Description

Astra DB is a serverless, cloud-native NoSQL database built on Apache Cassandra, providing vector search, multi-model support (tabular, search, graph), and elastic scalability for AI workloads with near-zero latency. Named a Forrester Leader for delivering NoSQL vector search capabilities on cloud.

2Hyper-converged Database (HCD)
CategoryCore product (on-premises/private cloud NoSQL database)
Description

HCD is the on-premises or private cloud counterpart to Astra DB, delivering the same vector-enabled NoSQL database capabilities for organizations that require on-premise deployment. Delivered as DataStax with IBM watsonx.data Premium edition; available in HCD 2.0, 1.2, and 1.1 versions.

3Langflow
CategoryOpen-source developer tool / low-code IDE
Description

Langflow is an open-source low-code visual development environment (100,000+ GitHub stars) for prototyping, building, and deploying retrieval-augmented generation and multi-agent AI applications. Built in Python, designed to work across models, APIs, and databases, and integrating with IBM watsonx Orchestrate as middleware.

4Astra Streaming
CategoryManaged streaming service
Description

Astra Streaming is a fully managed event streaming service based on Apache Pulsar that enables efficient data streaming, real-time processing, and CDC for Astra DB Serverless.

5Astra CLI
CategoryCommand-line tool
Description

Astra CLI is a command-line interface (v1.0.0 released November 2025) for managing Astra resources — databases, streaming tenants, keyspaces, tokens, organizations — through scripts or the local terminal, with colorized output, extended auto-completions, and Windows support.

6Mission Control
CategoryCluster management platform
Description

Mission Control is a Kubernetes-based platform for managing and orchestrating DSE and Cassandra clusters with declarative resources (MissionControlCluster, CassandraDatacenter), lifecycle manager (LCM) integration, and bundled observability components.

7DataStax AI Platform
CategoryAI Platform as a Service (PaaS)
Description

The DataStax AI Platform, built with NVIDIA AI (including NVIDIA NeMo Retriever and NIM microservices), is a platform-as-a-service for enterprise generative AI development that integrates Astra DB vector search with NVIDIA foundation model tooling.

8DataStax Enterprise (DSE)
CategoryEnterprise database platform
Description

DSE is the on-premises enterprise database platform based on Apache Cassandra with Advanced Workloads including DSE Search (Apache Solr), DSE Analytics (Apache Spark), and DSE Graph.

9OpsCenter
CategoryDatabase operations and monitoring tool
Description

OpsCenter (version 6.8) is a browser-based operations and monitoring tool for DataStax Enterprise (DSE) clusters providing backup/restore, repair, configuration management, and cluster visualization.

10DataStax Bulk Loader (DSBulk)
CategoryData loading utility
Description

DSBulk is a high-performance bulk loading and unloading utility for Cassandra-based databases supporting CSV, JSON, and other formats for large-scale data migrations and ETL.

11DataStax Studio
CategoryDeveloper IDE
Description

DataStax Studio is an interactive developer notebook IDE for working with CQL, Gremlin, and Cassandra-based databases, supporting query prototyping, visualization, and code execution.

12Stargate
CategoryOpen-source API gateway
Description

Stargate is an open-source data API gateway for Cassandra that provides Document, REST, GraphQL, and gRPC APIs on top of Cassandra databases.

13K8ssandra
CategoryKubernetes operator for Cassandra
Description

K8ssandra is an open-source Kubernetes operator and distribution for Apache Cassandra that provides automated operations, scaling, backup, and observability on Kubernetes clusters.

14CDC for Apache Cassandra
CategoryChange data capture
Description

CDC for Apache Cassandra provides change data capture functionality for Cassandra databases, capturing row-level changes in real time for streaming into downstream systems.

Scale indicator14 records

Each record includes

Type, Value, Description, Source

Partnership16 partners
Strategic tierFlagshipTypeStrategic or Co-development PartnerAnnounced on2025-02-26
Description

IBM acquired DataStax in February 2025 for approximately $2.1 billion to complement the watsonx portfolio. DataStax is now an IBM company, bringing Astra DB, HCD, and Langflow to watsonx. The acquisition strengthens IBM's enterprise AI data infrastructure capabilities for unstructured and multimodal data management.

Strategic tierCoreTypeStrategic or Co-development PartnerAnnounced on2024-12-01
Description

Wikimedia Deutschland launched an AI Knowledge Project in collaboration with DataStax, built with NVIDIA AI. The project integrated DataStax Astra DB on IBM watsonx.data to enhance access to Wikidata's multilingual knowledge graph, achieving a 30-fold increase in query speed and a 90% reduction in development time.

Strategic tierFlagshipTypeStrategic or Co-development PartnerAnnounced on2024-10-20
Description

DataStax and NVIDIA jointly built the DataStax AI Platform using NVIDIA AI technology. The platform integrates NVIDIA NIM microservices and NeMo Retriever for retrieval-augmented generation. Co-marketed launches have been conducted, including a project with Wikimedia Deutschland using DataStax Astra DB built with NVIDIA AI.

Strategic tierCoreTypeTechnology or IntegrationAnnounced on2024-09-27
Description

Jina AI collaborated with DataStax and Wikimedia Deutschland to launch a semantic search solution for non-profit AI developers. Jina AI is also listed as an integrated embedding provider within Astra DB Serverless (Astra vectorize).

Strategic tierCoreTypeTechnology or Integration
Description

OpenAI is integrated as a vectorize embedding provider within Astra DB Serverless, enabling automatic generation of embeddings for vector operations via Astra vectorize.

Strategic tierCoreTypeTechnology or Integration
Description

Hugging Face is integrated as an external embedding provider within Astra DB Serverless via Astra vectorize, available in both Dedicated and Serverless integration modes.

Strategic tierCoreTypeTechnology or Integration
Description

Azure OpenAI is integrated as a vectorize embedding provider within Astra DB Serverless for automated embedding generation.

Strategic tierCoreTypeTechnology or Integration
Description

LangChain is integrated with Astra DB Serverless and is the foundational framework for Langflow (DataStax's low-code AI builder). Available in Python and JavaScript variants.

Strategic tierCoreTypeTechnology or Integration
Description

Astra DB Serverless is available on AWS with integrations to AWS Bedrock, AWS SageMaker, AWS Glue, AWS Lambda, and AWS PrivateLink. Billing is supported via AWS Marketplace.

Strategic tierCoreTypeTechnology or Integration
Description

Astra DB Serverless is available on Google Cloud with integrations to Google Cloud Functions, Google Dataflow, Google Vertex AI, and Google Cloud Private Service Connect. Billing is supported via Google Cloud Marketplace.

Strategic tierCoreTypeTechnology or Integration
Description

Astra DB Serverless is available on Microsoft Azure with integrations to Azure Functions and Azure Private Link. Billing is supported via Azure Marketplace.

Strategic tierCoreTypeTechnology or Integration
Description

Amazon Bedrock is integrated with Astra DB Serverless for managed foundation model access.

Strategic tierCoreTypeTechnology or Integration
Description

Google Vertex AI is integrated with Astra DB Serverless as an Extension and for Search and chat capabilities.

Strategic tierCoreTypeTechnology or Integration
Description

DataStax delivered Glean and Unstructured integrations to the AI Platform at the RAG++ Event in NYC, enabling enterprise search and RAG capabilities.

Strategic tierCoreTypeTechnology or Integration
Description

Unstructured Serverless integration is available with Astra DB Serverless for unstructured data ingestion, delivered at the RAG++ Event in NYC.

Strategic tierCoreTypeTechnology or Integration
Description

Astra DB MCP server is available as a replacement for the deprecated Astra DB GitHub Copilot extension, enabling AI agent integrations with Astra DB.

Recent move7 records

Each record includes

Date, Type, Title, Description, Source

Expansion highlight6 records

Each record includes

Type, Description

Peers10 records
TypeDirect peer
Description

Pinecone is a managed vector database purpose-built for production RAG and semantic search workloads. It competes head-to-head with Astra DB in the vector database category and is named alongside DataStax in market commentary on the segment.

TypeDirect peer
Description

Weaviate is an open-source vector database with hybrid search and modular AI-native capabilities. It is a directly comparable peer to Astra DB on vector search, RAG, and AI application workloads.

TypeDirect peer
Description

Qdrant is an open-source vector similarity search engine written in Rust, focused on high-performance vector and hybrid retrieval for AI applications. It competes directly with Astra DB's vector search capabilities.

TypeDirect peer
Description

Zilliz is the commercial entity behind the open-source Milvus vector database. Both Milvus and Zilliz Cloud are explicitly named alongside DataStax in the vector database market commentary, making it a direct peer.

TypeDirect peer
Description

Chroma is an open-source embedding database widely used for LLM application development and prototyping. It overlaps with Astra DB on vector storage for AI applications and competes for the same AI/ML developer mindshare as Langflow-backed Astra workflows.

TypeBroad incumbent
Description

MongoDB is a leading general-purpose NoSQL document database that has added native vector search capabilities via Atlas Vector Search. It is a broader incumbent competing with Astra DB for enterprise document and AI workloads across hybrid cloud.

TypeBroad incumbent
Description

Couchbase is a distributed NoSQL document and key-value database with vector search capabilities targeting enterprise and edge workloads. Its distributed architecture and Cassandra-like operational profile make it a comparable broad incumbent to DataStax in enterprise NoSQL.

TypeBroad incumbent
Description

Elastic is a search and analytics platform with dense vector retrieval, hybrid search, and RAG capabilities integrated with Elasticsearch. It is a broader incumbent competing for enterprise search and AI infrastructure budgets alongside Astra DB.

TypeBroad incumbent
Description

Redis is an in-memory data store that has added vector search (RedisVL) and is positioned for low-latency AI application use cases. Its presence in real-time application stacks makes it a broad incumbent peer for DataStax's low-latency AI data infrastructure.

TypeEmerging player
Description

Confluent is the commercial provider of Apache Kafka and a streaming data platform, comparable to DataStax's Astra Streaming built on Apache Pulsar. It competes in the same managed event-streaming category, including for CDC and AI data pipelines.

Market position
Strengths5 records

Each record includes

Headline, Details, Source

Weaknesses5 records

Each record includes

Headline, Details, Source

Competitive moat6 records

Each record includes

Type, Details

Key risks6 records

Each record includes

Headline, Details, Source

Key highlights7 records

Each record includes

Headline, Details, Source

Customer concentration

Classification, Details

Named customers1 record

Each record includes

Name, Industry, Type, Use case, Source, UUID

Segment5 records

Each record includes

Title, Type, Primary, Description, Pain point addressed, Use case, Source

Ideal customer profile5 records

Each record includes

Profile, Firmographic size, Sales motion, Sales cycle length, Buying structure, Purchase trigger, Buyer persona, Geography, Industry vertical, Primary use case, Description, Pain points, Evidence proof points, Target buyer

Technology focused
Yes
API detail
Has APIbool
Yes

Docs URL, Description

Integration58 records

Each record includes

Title, Type, Description, Source

AI capability10 records

Each record includes

Type, Description, Source

AI maturity
App detail

Has app

Feature12 records

Each record includes

Title, Differentiator, Description, Source

Core technology
Revenue estimate
Valuation estimate
Number of profiles
Profiles14 records

Each record includes

Name, Designation, Designation category, Overview, Profile commentary, Source

Subsidiaries1 record

Each record includes

Name, Acquired on, Relationship type, Type, Business focus

Compliance4 records

Each record includes

Name, Class, Description

Funding overview

Funding stage, Last funding date, Total funding USD

Funding rounds10 records

Each record includes

Round, Amount USD, Date, Pre money valuation, Total investors, Investors, News

Investors23 records

Each record includes

Name, Type, Date of entry, Rounds participated, Website

Funding detail is available on the Subscription and Enterprise plan.Contact sales →

M&A6 records

Each record includes

Name, Acquisition type, Announced date, Completed date, Status, Website, News

Investment

Each record includes

Name, Round, Announced date, Lead investor, Website, News

M&A and investment is available on the Subscription and Enterprise plan.Contact sales →

DataStax

Cloud Database (Vector/NoSQL)datastax.com

DataStax provides a real-time data platform for generative AI applications, centered on Astra DB (a serverless Cassandra-based vector database), Langflow (an open-source low-code AI builder), and streaming products. It serves enterprise, digital-native, and platform-engineering teams, and now operates as an IBM company following a February 2025 acquisition.

What DataStax does

DataStax builds a real-time data platform for generative AI applications, headquartered in Santa Clara, California and founded in 2010. Its product portfolio centers on Astra DB, a serverless NoSQL and vector database built on Apache Cassandra; Hyper-converged Database (HCD), the on-premises counterpart; Astra Streaming, a managed Apache Pulsar service; and the DataStax Enterprise (DSE) data layer. The company acquired Langflow in 2024, an open-source low-code builder for AI applications and multi-agent workflows that has surpassed 100,000 GitHub stars. In October 2024 DataStax launched its AI Platform built with NVIDIA AI, integrating NVIDIA NIM and NeMo Retriever for retrieval-augmented generation. The platform holds HIPAA, SOC 2, ISO 27001, and PCI DSS certifications, and has been recognized as a Leader by both Forrester and Gartner in vector and cloud-native database evaluations.

DataStax's go-to-market combines a freemium developer tier (Astra DB) with enterprise subscriptions for Astra DB, DSE, and HCD, usage-based consumption via cloud marketplaces (AWS, Azure, GCP), and professional services. Pricing is anchored on Astra DB plan tiers, DSE/HCD node-based licensing, and per-capacity-unit (PCU) metering for streaming and cloud consumption. The customer base spans regulated enterprises, digital-native builders, and platform-engineering teams running Gen AI workloads, with Wikimedia Deutschland as a publicly cited production deployment. The company raised approximately $115 million from Goldman Sachs in 2022 at a $1.6 billion valuation and was acquired by IBM in February 2025 for roughly $2.1 billion, after which it operates as "DataStax, an IBM company." The employee count is reported in the 501–1,000 range, distributed across the U.S., EMEA, and Asia following the 2022 China expansion.

DataStax firmographics

Firmographics
Name
DataStax
Legal name
DataStax Inc.
Website
https://datastax.com
Company type
Private
Founded year
2010
Operating status
Acquired
Headcount range
501–1,000 employees
Short description
DataStax provides a real-time data platform for generative AI applications, centered on Astra DB (a serverless Cassandra-based vector database), Langflow (an open-source low-code AI builder), and streaming products. It serves enterprise, digital-native, and platform-engineering teams, and now operates as an IBM company following a February 2025 acquisition.
Ownership category
akta.pro rank

DataStax industry classification

Industry
Product category
Cloud Database (Vector/NoSQL)
NAICS
Custom Computer Programming Services (541511), Software Publishers (5132), Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (5182)
SIC
Services-Computer Programming, Data Processing, Etc. (7370), Services-Prepackaged Software (7372)
akta.pro primary industry
Enterprise AI Data & Knowledge Platforms (Vector Databases, Knowledge Graphs) (HDAEANAH)
akta.pro secondary industries
Managed Databases (Relational/NoSQL/In-Memory DBaaS) (HDABAAAE), Database-as-a-Service (DBaaS) Platforms (HDAEAAAN), Database Tools & Ecosystem (Replication, Backup/Recovery, HA/DR, Monitoring) (HDAEAAAO), Data Sharing, Data Exchange & Data Marketplace Platforms (HDAEABAH), Data Platform (Unified Data & Analytics) Suites (HDAEABAD)

Keywords

  • Vector database
  • NoSQL database
  • Retrieval-augmented generation
  • AI data platform
  • Generative AI infrastructure

Where DataStax is headquartered

Location

Headquarters

HQ city
Santa Clara
HQ country
United States
HQ region
North America

Offices6 records

Markets served

DataStax business model

Business model
GTM type
B2B
Offering type
Software
Cost components
Personnel, Technology or R&D, Infrastructure, Marketing or Sales, Operations

Revenue model

  1. Astra DB Subscription Plans: Recurring subscription revenue from the Astra DB Serverless product with multiple tiers including Free, Standard (Pay-As-You-Go and Subscription), and Enterprise plans. The Marketplace plan (now Standard) is billed through cloud provider marketplaces.
  2. Provisioned Capacity Units (PCUs): Usage-based revenue from Provisioned Capacity Units (PCUs) with Reserved Capacity Units (RCUs) for committed workloads and Hourly Capacity Units (HCUs) for flexible/auto-scaled capacity. PCUs support complex, latency-sensitive enterprise workloads with dedicated tenant environments.
  3. DataStax Enterprise (DSE) & Hyper-converged Database (HCD) Licenses: On-premises and private cloud database software licensing for DSE and HCD, including support subscriptions, with DataStax Enterprise Premium edition delivered as DataStax with IBM watsonx.data Premium edition.
  4. Enterprise Support Services & Professional Services: Premium support plans, consulting, and integration services bundled with enterprise contracts, including contact support, knowledge base access, and DataStax documentation services.
  5. Cloud Marketplace Consumption: Consumption-based revenue transacted through AWS, Google Cloud, and Microsoft Azure marketplaces where customers pay via the relevant cloud provider for their Astra DB usage.

Pricing tiers

ModelBillingPrice
FreemiumMonthlyFree plan with monthly credits
Usage-basedPay-as-you-goStandard plan (formerly Marketplace)
SubscriptionAnnualEnterprise plan
HybridPay-as-you-goProvisioned Capacity Units (PCUs)
FreemiumMonthlyFree trial for Astra DB

Go-to-market motion5 records

Distribution channels5 records

Marketing channels10 records

DataStax product offering

Product offering

Core offering

DataStax provides a serverless NoSQL and vector database (Astra DB) built on Apache Cassandra, with an on-premises counterpart (Hyper-converged Database, HCD) and an open-source low-code AI builder (Langflow) for prototyping, building, and deploying retrieval-augmented generation and multi-agent AI applications. The DataStax AI Platform, built with NVIDIA AI, integrates these databases with NVIDIA NeMo Retriever and NIM microservices to support enterprise generative AI workloads across on-prem, hybrid, and multi-cloud environments.

Product overview

DataStax, an IBM company, offers a unified AI-ready data platform for managing real-time, unstructured, and multimodal data at scale. The core of the platform is Astra DB (serverless NoSQL vector database built on Apache Cassandra) and HCD (Hyper-converged Database, the on-premises/private cloud counterpart), which together provide vector search, multi-model support (tabular, search, graph), and elastic scalability with near-zero latency. Complementing these are Langflow (an open-source low-code IDE for prototyping, building, and deploying RAG and multi-agent AI applications, with 100,000+ GitHub stars) and the DataStax AI Platform (built with NVIDIA AI, including NeMo Retriever and NIM microservices). Supporting products include Astra Streaming (managed Apache Pulsar event streaming), Mission Control (Kubernetes-based cluster management), Astra CLI (resource management CLI), DataStax Enterprise (DSE) on-premises database platform, OpsCenter (monitoring), DataStax Bulk Loader (DSBulk), DataStax Studio (interactive IDE), Stargate (open-source API gateway), and K8ssandra (Kubernetes operator). As part of IBM watsonx, DataStax delivers vector search, retrieval-augmented generation, and unstructured data management for enterprise AI applications running on-prem, hybrid, or multi-cloud.

Differentiator

Problem solved

Functional benefit

Brands

  • Astra DB: DataStax's serverless, multi-cloud NoSQL database built on Apache Cassandra, with native vector search capabilities and integrations for generative AI workloads.
  • Hyper-converged Database (HCD)
  • Langflow
  • Astra Streaming
  • DataStax Enterprise (DSE)
  • Mission Control

Products and services

  • Astra DB Astra DB is a serverless, cloud-native NoSQL database built on Apache Cassandra, providing vector search, multi-model support (tabular, search, graph), and elastic scalability for AI workloads with near-zero latency. Named a Forrester Leader for delivering NoSQL vector search capabilities on cloud.
  • Hyper-converged Database (HCD) HCD is the on-premises or private cloud counterpart to Astra DB, delivering the same vector-enabled NoSQL database capabilities for organizations that require on-premise deployment. Delivered as DataStax with IBM watsonx.data Premium edition; available in HCD 2.0, 1.2, and 1.1 versions.
  • Langflow Langflow is an open-source low-code visual development environment (100,000+ GitHub stars) for prototyping, building, and deploying retrieval-augmented generation and multi-agent AI applications. Built in Python, designed to work across models, APIs, and databases, and integrating with IBM watsonx Orchestrate as middleware.
  • Astra Streaming Astra Streaming is a fully managed event streaming service based on Apache Pulsar that enables efficient data streaming, real-time processing, and CDC for Astra DB Serverless.
  • Astra CLI Astra CLI is a command-line interface (v1.0.0 released November 2025) for managing Astra resources — databases, streaming tenants, keyspaces, tokens, organizations — through scripts or the local terminal, with colorized output, extended auto-completions, and Windows support.
  • Mission Control Mission Control is a Kubernetes-based platform for managing and orchestrating DSE and Cassandra clusters with declarative resources (MissionControlCluster, CassandraDatacenter), lifecycle manager (LCM) integration, and bundled observability components.
  • DataStax AI Platform The DataStax AI Platform, built with NVIDIA AI (including NVIDIA NeMo Retriever and NIM microservices), is a platform-as-a-service for enterprise generative AI development that integrates Astra DB vector search with NVIDIA foundation model tooling.
  • DataStax Enterprise (DSE) DSE is the on-premises enterprise database platform based on Apache Cassandra with Advanced Workloads including DSE Search (Apache Solr), DSE Analytics (Apache Spark), and DSE Graph.
  • OpsCenter OpsCenter (version 6.8) is a browser-based operations and monitoring tool for DataStax Enterprise (DSE) clusters providing backup/restore, repair, configuration management, and cluster visualization.
  • DataStax Bulk Loader (DSBulk) DSBulk is a high-performance bulk loading and unloading utility for Cassandra-based databases supporting CSV, JSON, and other formats for large-scale data migrations and ETL.
  • DataStax Studio DataStax Studio is an interactive developer notebook IDE for working with CQL, Gremlin, and Cassandra-based databases, supporting query prototyping, visualization, and code execution.
  • Stargate Stargate is an open-source data API gateway for Cassandra that provides Document, REST, GraphQL, and gRPC APIs on top of Cassandra databases.
  • K8ssandra K8ssandra is an open-source Kubernetes operator and distribution for Apache Cassandra that provides automated operations, scaling, backup, and observability on Kubernetes clusters.
  • CDC for Apache Cassandra CDC for Apache Cassandra provides change data capture functionality for Cassandra databases, capturing row-level changes in real time for streaming into downstream systems.

Quantifiable outcome

  • 30-fold increase in query speed and 90% reduction in development time for Wikimedia Deutschland's Wikidata multilingual knowledge graph
  • +3 more outcomes

Companies that use DataStax

Customer profile

Named customers1 record

Segments5 records

Ideal customer profiles5 records

DataStax technology and API

Technology

Technology focussed Yes

API detail

Has API
Yes
API docs
API detail

Core technology

AI maturity

App detail

Integration58 records

AI capability10 records

Feature12 records

DataStax partnerships and signals

Strategic signal

Partnerships

16 partnerships are on record, tiered flagship and core.

  • IBM (International Business Machines)flagshipStrategic or Co-development Partner · 26 February 2025IBM acquired DataStax in February 2025 for approximately $2.1 billion to complement the watsonx portfolio. DataStax is now an IBM company, bringing Astra DB, HCD, and Langflow to watsonx. The acquisition strengthens IBM's enterprise AI data infrastructure capabilities for unstructured and multimodal data management.
  • Wikimedia DeutschlandcoreStrategic or Co-development Partner · 1 December 2024Wikimedia Deutschland launched an AI Knowledge Project in collaboration with DataStax, built with NVIDIA AI. The project integrated DataStax Astra DB on IBM watsonx.data to enhance access to Wikidata's multilingual knowledge graph, achieving a 30-fold increase in query speed and a 90% reduction in development time.
  • NVIDIAflagshipStrategic or Co-development Partner · 20 October 2024DataStax and NVIDIA jointly built the DataStax AI Platform using NVIDIA AI technology. The platform integrates NVIDIA NIM microservices and NeMo Retriever for retrieval-augmented generation. Co-marketed launches have been conducted, including a project with Wikimedia Deutschland using DataStax Astra DB built with NVIDIA AI.
  • Jina AIcoreTechnology or Integration · 27 September 2024Jina AI collaborated with DataStax and Wikimedia Deutschland to launch a semantic search solution for non-profit AI developers. Jina AI is also listed as an integrated embedding provider within Astra DB Serverless (Astra vectorize).
  • OpenAIcoreTechnology or IntegrationOpenAI is integrated as a vectorize embedding provider within Astra DB Serverless, enabling automatic generation of embeddings for vector operations via Astra vectorize.
  • Hugging FacecoreTechnology or IntegrationHugging Face is integrated as an external embedding provider within Astra DB Serverless via Astra vectorize, available in both Dedicated and Serverless integration modes.
  • Microsoft Azure OpenAIcoreTechnology or IntegrationAzure OpenAI is integrated as a vectorize embedding provider within Astra DB Serverless for automated embedding generation.
  • LangChaincoreTechnology or IntegrationLangChain is integrated with Astra DB Serverless and is the foundational framework for Langflow (DataStax's low-code AI builder). Available in Python and JavaScript variants.
  • Amazon Web Services (AWS)coreTechnology or IntegrationAstra DB Serverless is available on AWS with integrations to AWS Bedrock, AWS SageMaker, AWS Glue, AWS Lambda, and AWS PrivateLink. Billing is supported via AWS Marketplace.
  • Google Cloud PlatformcoreTechnology or IntegrationAstra DB Serverless is available on Google Cloud with integrations to Google Cloud Functions, Google Dataflow, Google Vertex AI, and Google Cloud Private Service Connect. Billing is supported via Google Cloud Marketplace.
  • Microsoft AzurecoreTechnology or IntegrationAstra DB Serverless is available on Microsoft Azure with integrations to Azure Functions and Azure Private Link. Billing is supported via Azure Marketplace.
  • Amazon BedrockcoreTechnology or IntegrationAmazon Bedrock is integrated with Astra DB Serverless for managed foundation model access.
  • Google Vertex AIcoreTechnology or IntegrationGoogle Vertex AI is integrated with Astra DB Serverless as an Extension and for Search and chat capabilities.
  • GleancoreTechnology or IntegrationDataStax delivered Glean and Unstructured integrations to the AI Platform at the RAG++ Event in NYC, enabling enterprise search and RAG capabilities.
  • Unstructured.iocoreTechnology or IntegrationUnstructured Serverless integration is available with Astra DB Serverless for unstructured data ingestion, delivered at the RAG++ Event in NYC.
  • Model Context Protocol (MCP)coreTechnology or IntegrationAstra DB MCP server is available as a replacement for the deprecated Astra DB GitHub Copilot extension, enabling AI agent integrations with Astra DB.

Scale indicators14 records

Recent moves7 records

Expansion highlights6 records

DataStax competitors and assessment

Company assessment

Direct peers

  • Pinecone: Pinecone is a managed vector database purpose-built for production RAG and semantic search workloads. It competes head-to-head with Astra DB in the vector database category and is named alongside DataStax in market commentary on the segment.
  • Weaviate: Weaviate is an open-source vector database with hybrid search and modular AI-native capabilities. It is a directly comparable peer to Astra DB on vector search, RAG, and AI application workloads.
  • Qdrant: Qdrant is an open-source vector similarity search engine written in Rust, focused on high-performance vector and hybrid retrieval for AI applications. It competes directly with Astra DB's vector search capabilities.
  • Zilliz (Milvus): Zilliz is the commercial entity behind the open-source Milvus vector database. Both Milvus and Zilliz Cloud are explicitly named alongside DataStax in the vector database market commentary, making it a direct peer.
  • Chroma: Chroma is an open-source embedding database widely used for LLM application development and prototyping. It overlaps with Astra DB on vector storage for AI applications and competes for the same AI/ML developer mindshare as Langflow-backed Astra workflows.

Broad incumbents

  • MongoDB: MongoDB is a leading general-purpose NoSQL document database that has added native vector search capabilities via Atlas Vector Search. It is a broader incumbent competing with Astra DB for enterprise document and AI workloads across hybrid cloud.
  • Couchbase: Couchbase is a distributed NoSQL document and key-value database with vector search capabilities targeting enterprise and edge workloads. Its distributed architecture and Cassandra-like operational profile make it a comparable broad incumbent to DataStax in enterprise NoSQL.
  • Elastic: Elastic is a search and analytics platform with dense vector retrieval, hybrid search, and RAG capabilities integrated with Elasticsearch. It is a broader incumbent competing for enterprise search and AI infrastructure budgets alongside Astra DB.
  • Redis: Redis is an in-memory data store that has added vector search (RedisVL) and is positioned for low-latency AI application use cases. Its presence in real-time application stacks makes it a broad incumbent peer for DataStax's low-latency AI data infrastructure.

Emerging players

  • Confluent: Confluent is the commercial provider of Apache Kafka and a streaming data platform, comparable to DataStax's Astra Streaming built on Apache Pulsar. It competes in the same managed event-streaming category, including for CDC and AI data pipelines.

Market position

Strengths5 records

Weaknesses5 records

Competitive moat6 records

Key risks6 records

Key highlights7 records

Customer concentration

DataStax social profiles

Digital presence

DataStax compliance and trust

Trust signal

Compliance4 records

DataStax financial estimates

Financial estimate

Revenue estimate

Valuation estimate

DataStax leadership team

Management profile

Number of profiles

Profiles14 records

DataStax subsidiaries and ownership

Company hierarchy

Subsidiaries1 record

DataStax funding detail

Funding detail

Funding overview

Funding rounds10 records

Investors23 records

Funding detail is available on the Subscription and Enterprise plan.Contact sales →

DataStax M&A and investment

M&A and investment

M&A6 records

Investments

M&A and investment is available on the Subscription and Enterprise plan.Contact sales →

Frequently asked questions about DataStax

What does DataStax do?

DataStax provides a serverless NoSQL and vector database (Astra DB) built on Apache Cassandra, with an on-premises counterpart (Hyper-converged Database, HCD) and an open-source low-code AI builder (Langflow) for prototyping, building, and deploying retrieval-augmented generation and multi-agent AI applications. The DataStax AI Platform, built with NVIDIA AI, integrates these databases with NVIDIA NeMo Retriever and NIM microservices to support enterprise generative AI workloads across on-prem, hybrid, and multi-cloud environments.

Is DataStax a public or private company?

DataStax is a private company. It is classified as corporate owned and is currently acquired.

When was DataStax founded?

DataStax was founded in 2010. It employs 501 to 1,000 people.

Where is DataStax based?

DataStax is headquartered in Santa Clara, United States, in the North America region.

How does DataStax make money?

Five revenue lines are on record. Astra DB Subscription Plans are the primary driver. The others are provisioned Capacity Units (PCUs), dataStax Enterprise (DSE) & Hyper-converged Database (HCD) Licenses, enterprise Support Services & Professional Services and cloud Marketplace Consumption.

Who are DataStax's main competitors?

Direct peers on record are Pinecone, Weaviate, Qdrant, Zilliz (Milvus) and Chroma. Broad incumbents are MongoDB, Couchbase, Elastic and Redis. Confluent is listed as an emerging player.

Does DataStax have an API?

Yes. DataStax offers multiple APIs including the Data API (schema-less, document-based, JSON API for Astra DB Serverless vector databases with Python, TypeScript, Java, and C# client libraries) and the DevOps API (administrative REST API for managing Astra DB resources, organizations, databases, keyspaces, access lists, and metrics). The Data API enables developers to build production generative AI and retrieval-augmented generation (RAG) applications by inserting, finding, updating, and deleting documents in collections and rows in tables. The DevOps API v2 provides programmatic access to provision and manage Astra DB Serverless databases. Authentication is handled via application tokens or SSO. Clients support table-based structured data (in preview) and hybrid search. Developer documentation is at docs.datastax.com/en/astra-db-serverless/api-reference/dataapiclient.html.

What industry is DataStax in?

DataStax's product category is Cloud Database (Vector/NoSQL). Its primary akta.pro industry code is HDAEANAH, Enterprise AI Data & Knowledge Platforms (Vector Databases, Knowledge Graphs), with a secondary code of HDABAAAE, Managed Databases (Relational/NoSQL/In-Memory DBaaS). Its NAICS code is 541511 and its SIC code is 7370.

Unlock the full company data

50 free credits on sign-up, no credit card required.

Contact sales
Live signals
TechBullionNoSQL Databases Explained: What It Means for Consumers and Businesses in the USAThe NoSQL database market is projected to grow from $15.04 billion in 2025 to $69.09 billion by 2031, driven by demand for flexible data handling in finance and enterprise sectors. IBM's acquisition of NoSQL specialist DataStax signals a strategic shift toward integrating flexible data stores into core platform offerings.SD TimesModern Data & Knowledge Platforms: The Foundation Every AI Strategy Actually Runs On: SD Times 100The SD Times 100 2026 category Modern Data & Knowledge Platforms highlights companies like Cockroach Labs, Confluent, and Databricks for data infrastructure supporting AI. It notes that data architecture decisions are costly to unwind and that retrieval quality is now a product quality issue. The category also lists new 2026 additions such as LanceDB and MindsDB.openPR.comFuture Perspectives: Key Trends Shaping the Graph Database Vector Search Market Up to 2030The graph database vector search market is projected to reach $8.44 billion by 2030, growing at a compound annual growth rate of 23.3%, driven by advances in retrieval augmented generation, graph analytics, and real-time contextual search capabilities. In February 2025, IBM acquired DataStax Inc., a US data management firm specializing in graph database vector search technologies, integrating DataStax's NoSQL and vector database platforms including AstraDB and DataStax Enterprise. Neo4j Inc. introduced native vector search features in August 2023, enabling developers to combine vector similarity searches with graph relationship data to improve AI output accuracy.openPR.comWorldwide Trends Overview: The Rapid Evolution of the Database Management Services MarketThe Business Research Company published a market report showing the database management services market is projected to reach $45.02 billion by 2030, growing at a CAGR of 11.2%, driven by cloud migration, real-time database observability, and multi-cloud complexity. In February 2025, IBM acquired DataStax Inc. to accelerate its enterprise AI initiatives by integrating NoSQL and vector database technologies for generative AI applications. Additionally, Acceldata Inc. launched the industry's first agentic data management platform that autonomously manages enterprise data across hybrid and multi-cloud environments.openPR.comFuture Perspectives: Key Trends Shaping the Big Data as a Service (BDaaS) Market Until 2030IBM Corporation acquired DataStax Inc. in February 2025 to strengthen its artificial intelligence and data management capabilities, specifically enhancing support for unstructured data and generative AI technologies for enterprise customers. The broader BDaaS market is projected to grow to $116.25 billion by 2030, with a compound annual growth rate of 25.9%, driven by AI-driven analytics adoption, hybrid cloud strategies, and increasing demand for real-time decision-making solutions. Additionally, FactSet Research Systems Inc. launched a data-as-a-service solution for the financial sector in October 2024, designed to streamline data management and reduce operational costs.openPR.comGraph Database Market is expected to Hit US$ 12.8 billion by 2031 | Major Companies - Amazon Web Services, Inc., Cloud Software Group, Inc., IBM Corporation, Oracle, DataStax, MicrosoftDataM Intelligence released a market research report projecting the global Graph Database Market to grow from US$ 2.9 billion in 2023 to US$ 12.8 billion by 2031, with a CAGR of 20.2% during 2024-2031, driven by rising data complexity and AI-driven analytics demand. In notable M&A activity within the sector, TigerGraph acquired a data integration/graph processing startup in March 2026 to strengthen cloud-native graph analytics, while Neo4j acquired a niche graph analytics startup in February 2026 to expand GenAI integration capabilities. North America leads the market due to advanced infrastructure and early technology adoption, with Asia-Pacific identified as the fastest-growing region.Trend MicroVoid Dokkaebi Uses Fake Job Interview Lure to Spread Malware via Code RepositoriesTrendAI Research documented Void Dokkaebi's evolution from single-target social engineering into a self-propagating supply chain threat that turns compromised developer repositories into malware delivery channels. The campaign exploits VS Code workspace configurations and performs commit tampering to spread infection through trusted development workflows. Analysis in March 2026 identified over 750 infected repositories, more than 500 malicious VS Code task configurations, and 101 instances of a commit tampering tool, with organizational repositories from DataStax and Neutralinojs confirmed as compromised.Trend MicroVoid Dokkaebi Uses Fake Job Interview Lure to Spread Malware via Code RepositoriesTrendAI Research published an analysis in March 2026 detailing Void Dokkaebi's evolution from single-target social engineering into a self-propagating supply chain threat that exploits fake job interviews to compromise developers and weaponize their code repositories. The campaign, attributed to a North Korea-aligned threat actor, has infected over 750 repositories and uses blockchain infrastructure (Tron, Aptos, Binance Smart Chain) for payload staging beyond traditional takedowns. Organizations including DataStax and Neutralinojs were identified as victims, with repositories carrying infection markers that propagate the malware to downstream developers and collaborators.The New StackWhat engineering leaders get wrong about data stack consolidationAnil Inamdar argues that data stack consolidation creates architectural debt, as proprietary layers reduce portability and engineer expertise. He warns that open-source tools are no longer neutral, and leaders must evaluate migration feasibility. The trend of consolidation will continue, forcing deliberate design choices.Market Research FutureGraph Technology Market Size, Share, Growth Report 2035Market Research Future has published a comprehensive report on the Graph Technology Market, estimating its 2024 valuation at $4.567 billion and projecting growth to $20.81 billion by 2035, representing a compound annual growth rate (CAGR) of 14.78% during the forecast period 2025–2035. The report identifies key market drivers including the rising demand for data connectivity, advancements in graph algorithms, integration with artificial intelligence, and the growing emphasis on real-time analytics across sectors such as finance, healthcare, and telecommunications. North America leads the market with approximately 45% global share, while Asia-Pacific emerges as the fastest-growing region; key players profiled include Neo4j, Amazon Web Services, Microsoft, IBM, Oracle, TigerGraph, DataStax, ArangoDB, SAP, and Cytoscape.