YData
YData provides a data-centric AI platform combining synthetic data generation, automated data profiling, and data preparation pipelines for ML development. Its YData Fabric enterprise platform and open-source SDK serve enterprises and data scientists across financial services, telecom, healthcare, and retail.
- Company typePrivate
- Founded2018
- HeadquartersSeattle, United States
- Headcount11–50
- GTM typeB2B
- OfferingSoftware
What YData does
YData is a data-centric AI platform company that builds tools for improving the quality, privacy, and utility of data used in machine learning. Founded in 2018 in Lisbon, Portugal, with current headquarters in Seattle, the company develops two principal product lines: YData Fabric, a Kubernetes-native enterprise platform combining an interactive Data Catalog, generative AI-powered synthetic data generation, and automated data preparation pipelines; and the YData SDK, an open-source Python toolkit (available on PyPi) for data profiling and synthetic data generation. Underlying technology relies on generative AI techniques including GANs to produce synthetic tabular, time-series, and full-database datasets that mimic real-data statistical properties while preserving privacy, supported by automated quality and privacy controls using divergence metrics, correlation measures, non-parametric tests, TSTR methodology, and inference-attack testing.
The company's flagship open-source project, ydata-profiling (formerly Pandas Profiling), has accumulated 52M+ downloads and 10,000+ GitHub stars, with 12,000+ data scientists reportedly using YData products daily. YData sells through a product-led growth motion with a self-serve dashboard (dashboard.ydata.ai), a freemium SDK, monthly and annual subscriptions for the Fabric Platform, and enterprise sales complemented by Azure Marketplace and AWS Marketplace listings and on-premises Kubernetes deployment. Named enterprise customers include EDP Distribuição, Revolut, Telefonica, NextBrain.ai, Lovys, and Ciclo Mobility, with primary verticals being Financial Services, Telecommunications, Healthcare, and Retail; secondary use cases exist in Utilities.
The company raised approximately $3.24M in disclosed funding (a $2.7M seed round in October 2021 led by Flying Fish Partners, alongside earlier Techstars, Faber, and European Data Incubator support) and joined a ~€80M Responsible AI consortium in November 2022 as the only data-focused member. In October 2025, KPMG LLP acquired YData Labs Inc.'s intellectual property and technology assets to establish a synthetic data center of excellence, representing a significant corporate transaction that shapes the company's forward operating trajectory. The team is small (11-50 employees), the company remains privately held, and headquarters are in Seattle, Washington with an engineering/R&D presence in Lisbon, Portugal.
YData firmographics
Firmographics- Name
- YData
- Legal name
- YData Labs Inc.
- Website
- https://ydata.ai
- Company type
- Private
- Founded year
- 2018
- Operating status
- Operating
- Headcount range
- 11–50 employees
- Short description
- YData provides a data-centric AI platform combining synthetic data generation, automated data profiling, and data preparation pipelines for ML development. Its YData Fabric enterprise platform and open-source SDK serve enterprises and data scientists across financial services, telecom, healthcare, and retail.
- Ownership category
- akta.pro rank
YData industry classification
Industry- Product category
- Synthetic Data and Data Quality Platform
- NAICS
- Computer Systems Design and Related Services (54151), Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (518210)
- SIC
- Services-Prepackaged Software (7372), Services-Computer Programming Services (7371), Services-Computer Processing & Data Preparation (7374)
- akta.pro primary industry
- End-to-End MLOps & ML Platform Suites (HDAAABAA)
- akta.pro secondary industries
- Confidential AI & Privacy-Preserving ML (federated learning, MPC, HE, TEEs) (HDAAAKAI), MLOps/LLMOps & Model Lifecycle Management Services (BPAEAHAH)
Keywords
Where YData is headquartered
LocationHeadquarters
- HQ city
- Seattle
- HQ country
- United States
- HQ region
- North America
Offices2 records
Markets served
YData business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Personnel, Technology or R&D, Marketing or Sales, Operations
Revenue model
- YData Fabric Platform Subscription: Cloud-native platform deployed on customer infrastructure (Azure/AWS/on-premises), offering subscription-based access with monthly and annual billing options, including automatic renewal for subscriptions
- YData SDK/Package Subscription: Software development kit with pricing tiers including free trial, subscription-based access for data profiling and synthetic data generation capabilities
- Enterprise Professional Services: Enterprise deployments and support services potentially included in subscription tiers, with dedicated assistance for implementation
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Freemium | Monthly | Free Trial Available |
| Subscription | Annual | Fabric Platform - Enterprise Deployment |
| Subscription | Monthly | SDK/Package Pricing |
Go-to-market motion1 record
Distribution channels6 records
Marketing channels6 records
YData product offering
Product offeringCore offering
YData provides a data-centric AI platform that combines synthetic data generation, data profiling, and automated data preparation pipelines to improve data quality for machine learning and AI development. Its two flagship offerings are YData Fabric, a Kubernetes-native enterprise platform deployed on customer infrastructure (Azure, AWS, or on-premises), and YData SDK, an open-source Python package providing data profiling and synthetic data generation capabilities directly to data scientists.
Product overview
YData offers a dual-product portfolio consisting of YData Fabric Platform and YData SDK. YData Fabric is a kubernetes-native enterprise data quality platform combining three core modules: Data Catalog for data asset management and drift tracking, Synthetic Data for generative AI-powered data synthesis, and Pipelines for automated data preparation at scale. YData SDK is an open-source Python package providing data profiling and synthetic data generation capabilities directly to data scientists. The platform improves data quality for AI development, delivering up to 10x productivity gains, up to 25% faster AI model delivery, and up to 20% model performance boost through improved data quality.
Differentiator
Problem solved
Functional benefit
Brands
- YData Fabric: A comprehensive data-centric AI platform that provides data cataloging, synthetic data generation, and automated data pipelines for machine learning workflows.
- ydata-profiling
- YData SDK
Products and services
- YData Fabric Platform A Kubernetes-native enterprise data quality platform for AI development that combines data cataloging, synthetic data generation, and automated data preparation pipelines. Deployed on customer infrastructure (Azure, AWS, or on-premises) and targeted at enterprise data science and ML teams needing GDPR-compliant data workflows.
- YData SDK An open-source Python software development kit for data scientists providing data profiling and synthetic data generation capabilities. Available on PyPi with documentation at docs.sdk.ydata.ai, targeting individual data scientists and developers who want programmatic access to YData's data quality and synthetic data generation capabilities.
Quantifiable outcome
- 10x productivity improvement for data scientists
- +4 more outcomes
Companies that use YData
Customer profileNamed customers9 records
Segments6 records
Ideal customer profiles5 records
YData technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration3 records
AI capability3 records
Feature5 records
YData partnerships and signals
Strategic signalPartnerships
Four partnerships are on record, tiered flagship and core.
- KPMG LLPflagshipIn October 2025, KPMG LLP acquired the intellectual property and technology assets of YData Labs Inc. to establish a synthetic data center of excellence and provide privacy-compliant datasets for clients. This represents a significant validation of YData's technology and market positioning.
- Responsible AI ConsortiumflagshipYData participates in the world's largest artificial intelligence consortium for Responsible AI with an almost €80 million investment. YData is the only company in the consortium focused entirely on data, responsible for data understanding, democratization, bias mitigation, and data causality. Consortium includes visionary companies developing 21 AI products until 2030 and creating 210+ job opportunities.
- Microsoft AzurecoreYData Fabric is available in Azure Marketplace, enabling customers to deploy the platform directly in their Azure environment with simplified procurement and deployment.
- Amazon Web Services (AWS)coreYData Fabric is available in AWS Marketplace, allowing customers to deploy the platform in their AWS infrastructure.
Scale indicators6 records
Recent moves6 records
Expansion highlights6 records
YData competitors and assessment
Company assessmentDirect peers
- MOSTLY AI: MOSTLY AI is the closest direct competitor to YData — a synthetic data platform purpose-built for tabular and time-series data with privacy guarantees, serving financial services and regulated enterprises. It competes head-to-head with YData Fabric on benchmark-recognized synthetic data generation.
- Tonic.ai: Tonic.ai offers synthetic data and data masking for software development and ML training, with similar privacy-preserving positioning and enterprise go-to-market as YData. Directly overlaps on de-identification, test data, and ML training data use cases.
- Hazy: Hazy is an enterprise synthetic data vendor focused on financial services and regulated industries with comparable privacy-by-design architecture and benchmark-style positioning. Competes on the same procurement cycles as YData in banking and insurance.
- Gretel.ai: Gretel provides a synthetic data platform for ML and AI development with APIs and SDKs, including time-series and tabular support. Directly comparable in product surface area (generate, profile, transform) and developer-led distribution model.
- Datagen: Datagen generates synthetic data, primarily for computer vision, but has expanded into tabular ML use cases. Comparable in generative AI approach to data but with a different primary modality focus.
- Synthesized: Synthesized offers a data platform for synthetic data generation, data validation, and ML feature engineering aimed at regulated enterprises. Overlaps with YData's Fabric Platform in financial services and insurance verticals.
Emerging players
- Iguazio (acquired by McKinsey QuantumBlack): Iguazio is an MLOps platform focused on data and model pipelines for enterprise AI. Following its McKinsey/QuantumBlack acquisition, it now competes in the same enterprise data-for-AI space and follows a similar professional-services-led distribution model as YData's KPMG relationship.
Broad incumbents
- DataRobot: DataRobot is an established AI/ML platform vendor offering data preparation, feature engineering, and MLOps capabilities that overlap with YData's pipeline and profiling modules at enterprise scale.
- Amazon Web Services (SageMaker): AWS SageMaker provides synthetic data and data quality tooling as part of a broad ML platform — YData Fabric is already listed on the AWS Marketplace. Represents the bundling threat that could commoditize standalone synthetic data vendors.
- Microsoft Azure Machine Learning: Azure ML ships data prep, synthetic data, and Responsible AI features; YData Fabric is also a marketplace partner. The most significant bundling competitor for enterprise data science workflows.
Market position
Strengths4 records
Weaknesses4 records
Competitive moat5 records
Key risks5 records
Key highlights7 records
Customer concentration
YData social profiles
Digital presenceYData compliance and trust
Trust signalCompliance5 records
YData financial estimates
Financial estimateRevenue estimate
Valuation estimate
YData leadership team
Management profileNumber of profiles
Profiles4 records
YData funding detail
Funding detailFunding overview
Funding rounds5 records
Investors7 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
YData M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about YData
What does YData do?
YData provides a data-centric AI platform that combines synthetic data generation, data profiling, and automated data preparation pipelines to improve data quality for machine learning and AI development. Its two flagship offerings are YData Fabric, a Kubernetes-native enterprise platform deployed on customer infrastructure (Azure, AWS, or on-premises), and YData SDK, an open-source Python package providing data profiling and synthetic data generation capabilities directly to data scientists.
Is YData a public or private company?
YData is a private company. It is classified as venture growth investor backed and is currently operating.
When was YData founded?
YData was founded in 2018. It employs 11 to 50 people.
Where is YData based?
YData is headquartered in Seattle, United States, in the North America region.
How does YData make money?
Three revenue lines are on record. YData Fabric Platform Subscription is the primary driver. The others are YData SDK/Package Subscription and enterprise Professional Services.
Who are YData's main competitors?
Direct peers on record are MOSTLY AI, Tonic.ai, Hazy, Gretel.ai, Datagen and Synthesized. Iguazio (acquired by McKinsey QuantumBlack) is listed as an emerging player. Broad incumbents are DataRobot, Amazon Web Services (SageMaker) and Microsoft Azure Machine Learning.
Does YData have an API?
Yes. Users may access their data relating to this Application via the Application Program Interface (API). Documentation available at docs.sdk.ydata.ai. The YData SDK is available on PyPi at pypi.org/project/ydata-sdk/. Developer documentation is at docs.sdk.ydata.ai.
What industry is YData in?
YData's product category is Synthetic Data and Data Quality Platform. Its primary akta.pro industry code is HDAAABAA, End-to-End MLOps & ML Platform Suites, with a secondary code of HDAAAKAI, Confidential AI & Privacy-Preserving ML (federated learning, MPC, HE, TEEs). Its NAICS code is 54151 and its SIC code is 7372.