Datatang
Datatang is a Beijing-based AI training data services provider offering copyright datasets, custom data collection/annotation, vertical solutions, and the proprietary Shujiajia annotation platform, serving 5,000+ global enterprise customers across autonomous driving, LLM, speech, vision, and embodied AI.
- Company typePrivate
- Founded2010
- HeadquartersBeijing, China
- Headcount501–1,000
- GTM typeB2B
- OfferingSoftware
What Datatang does
Datatang (数据堂), founded in 2010 and headquartered in Beijing's Haidian District (Aobei Technology Park), is a privately held B2B AI training data services provider serving global enterprise customers building computer vision, speech, natural language, multi-modal, and large language model (LLM) systems. The company operates a full-stack service portfolio built around four pillars: (1) pre-built copyright datasets covering speech recognition/synthesis across 200+ languages and dialects (2M+ hours), computer vision (~1M IDs / 800TB), OCR, pronunciation dictionaries, and natural language understanding; (2) data customization services spanning multi-modal, LiDAR point cloud, street view, OCR, behavior/identity recognition, and speech; (3) vertical-specific industry solutions for autonomous driving, smart customer service, smart home, intelligent entertainment, new retail, and smart healthcare; and (4) the proprietary Shujiajia (数加加) data annotation platform supporting 2D, 3D, and 4D labeling with algorithm-assisted pre-recognition that delivers 30%+ efficiency improvements.
Underlying technology centers on the Shujiajia platform's algorithm-assisted pre-identification pipeline and an 8,000 sqm dedicated physical data collection factory for embodied AI/robotics research, complemented by an in-house professional annotation workforce and medical-specialist labeling teams. The company holds ISO9001, ISO27001, and ISO27701 certifications and secures national IP certificates on finished datasets, with a subject-authorized consent regime that addresses enterprise buyer due diligence.
Datatang's go-to-market is a sales-led enterprise motion: direct consultation via the website, customer service system, phone (13051623904), and email ([email protected]) under quote-based custom pricing on multi-year contracts. Revenue mechanics combine one-time copyright dataset licensing, professional services for custom collection/annotation, and recurring/subscription access to the Shujiajia platform and training edition. The company claims 5,000+ global enterprise customers served and a product footprint spanning China, Japan, Korea, the US, UK, France, Germany, and Russia.
Datatang firmographics
Firmographics- Name
- Datatang
- Legal name
- 数据堂(北京)科技股份有限公司
- Website
- https://datatang.com
- Company type
- Private
- Founded year
- 2010
- Operating status
- Operating
- Headcount range
- 501–1,000 employees
- Short description
- Datatang is a Beijing-based AI training data services provider offering copyright datasets, custom data collection/annotation, vertical solutions, and the proprietary Shujiajia annotation platform, serving 5,000+ global enterprise customers across autonomous driving, LLM, speech, vision, and embodied AI.
- Ownership category
- akta.pro rank
Datatang industry classification
Industry- Product category
- AI Training Data Services
- NAICS
- Custom Computer Programming Services (541511), Software Publishers (5132), Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (51821)
- SIC
- Services-Computer Programming, Data Processing, Etc. (7370), Services-Computer Processing & Data Preparation (7374)
- akta.pro primary industry
- Data Labeling & Annotation Services (HDAAALAB)
- akta.pro secondary industries
- Data Pipelines for GenAI (Curation, Filtering, Deduplication, Copyright) (HDAAACAJ), Dataset Discovery, Marketplaces & Licensing (HDAAALAE), Data Monetization & Data Products Services (BPAEAHAL), Data Augmentation & Programmatic Labeling (HDAAALAH)
Keywords
Where Datatang is headquartered
LocationHeadquarters
- HQ city
- Beijing
- HQ country
- China
- HQ region
- Asia
Offices1 record
Markets served
Datatang business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Personnel, Operations, Technology or R&D, Marketing or Sales, Infrastructure
Revenue model
- Copyright Dataset Sales: Sales of pre-built, copyright-cleared AI training datasets including speech, computer vision, NLP, and LLM data. All data is authorized by subjects, enabling immediate use by customers.
- Data Customization Services: Custom data collection and annotation services tailored to specific customer needs across various modalities including multi-modal, LiDAR point clouds, street view, OCR, behavior recognition, identity recognition, speech recognition, and speech synthesis.
- Industry Solutions: Vertical-specific AI training data solutions for industries including autonomous driving, intelligent customer service, smart home, intelligent entertainment, new retail, healthcare, high-quality dataset construction, and large model services.
- Data Annotation Platform Access: Access to the proprietary data annotation platform (数加加) for customer data labeling needs, including training platform for education purposes.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Other | Multi-year contract | Custom quotation-based pricing |
Go-to-market motion1 record
Distribution channels1 record
Marketing channels2 records
Datatang product offering
Product offeringCore offering
Datatang provides copyright-cleared AI training datasets, customized data collection and annotation services, vertical-specific industry solutions, and its self-developed Shujiajia annotation platform for buyers developing AI and machine learning models. Offerings span large language model, computer vision, speech recognition/synthesis, OCR, and embodied intelligence data, with subject authorization that lets customers deploy the data immediately.
Product overview
Datatang (数据堂) is a global AI training data service provider founded in 2010, offering a comprehensive full-stack data service portfolio. The core offerings include: (1) Training Datasets - copyright datasets for embodied AI, LLM, computer vision, speech recognition, speech synthesis, OCR, pronunciation dictionary, and natural language understanding; (2) Data Customization Services - tailored collection and processing for multimodal, LiDAR point cloud, street view, OCR, behavior recognition, identity recognition, speech recognition, and speech synthesis; (3) Industry Solutions - high-quality dataset construction, LLM solutions, and vertical solutions for smart driving, smart customer service, smart home, smart entertainment, new retail, and smart healthcare; (4) Data Annotation Platform (Shujiajia) - self-developed annotation tool supporting 2D/3D/4D labeling with algorithm-assisted automation. The company has built 1000+ copyright datasets, 2 million+ hours of speech data, and 800TB of computer vision data, covering 100+ languages and dialects, serving over 3000 global enterprise customers.
Differentiator
Problem solved
Functional benefit
Brands
- 数加加: Self-developed data annotation platform supporting 2D, 3D, and 4D data annotation for various AI training needs.
Products and services
- Embodied AI Training Datasets (具身智能训练数据集) Multi-scenario, real-world data collection for embodied AI applications, delivered as large-scale proprietary copyright finished datasets used to build high-precision, multi-modal intelligent robot training data loops.
- LLM Training Datasets (大模型训练数据集) PB-level copyrighted large language model datasets including large-scale high-quality unsupervised corpora, instruction-tuning Q&A pairs, and image-text-video multi-modal data for foundation model training.
- Computer Vision Training Datasets (计算机视觉训练数据集) Approximately 800TB of computer vision copyright datasets covering roughly 1 million IDs across face, body, traffic scenes, and multi-modal models for computer vision AI training.
- Speech Recognition Training Datasets (语音识别训练数据集) More than 10 million hours of speech datasets covering 200+ languages and dialects across varied recording equipment, scenarios, and formats for ASR model development.
- Speech Synthesis Training Datasets (语音合成训练数据集) Multi-language, multi-voice, multi-style speech synthesis resources recorded in NR15 acoustic environments with professional voice actors for TTS applications.
- OCR Training Datasets (OCR训练数据集) Optical character recognition training datasets supporting text extraction from images and documents across multiple languages and scenarios.
- Pronunciation Dictionary Training Datasets (发音词典训练数据集) Pronunciation dictionary datasets engineered for speech recognition and synthesis model training.
- Natural Language Understanding Training Datasets (自然语言理解训练数据集) Natural language understanding datasets for NLP model training, covering text classification, entity recognition, and semantic understanding tasks.
- Shujiajia Data Annotation Platform (数据标注平台 / 数加加) Self-developed annotation platform supporting 2D, 3D, and 4D data labeling with algorithm-assisted pre-identification for semi-automatic annotation, improving per-capita efficiency by over 30%.
- Data Annotation Training Platform (数据标注实训平台) Training platform for data annotation skill development and education, extending the Shujiajia platform for workforce training purposes.
- Multimodal Data Customization (多模态数据定制) Custom multimodal data collection and processing services tailored to specific AI training requirements.
- LiDAR Point Cloud Data Customization (激光雷达点云数据定制) Custom LiDAR point cloud data collection and annotation tailored for autonomous driving applications.
- High-Quality Dataset Construction Solution (高质量数据集建设解决方案) End-to-end industry dataset construction solution helping enterprises release data element value and continuously deliver high-quality, trustworthy training data for AI applications.
- LLM Solution (大模型解决方案) One-stop large model data services combining copyrighted datasets, model fine-tuning annotation, and model evaluation, delivered with proprietary full-stack LLM annotation tools.
- Smart Driving Solution (智能驾驶解决方案) Autonomous driving data solution offering large-scale proprietary finished datasets covering in-cabin intelligence R&D and supporting 2D, 3D, and 4D annotation for full driving coverage.
- Smart Customer Service Solution (智能客服解决方案) Customer service AI solution leveraging 2 million+ hours of speech data to improve voice recognition models and multi-type data processing for AI model performance.
- Smart Home Solution (智能家居解决方案) Smart home data solution covering multi-scenario, multi-modal, and diverse data types to improve sensing and interaction precision for home devices.
- Smart Entertainment Solution (智能娱乐解决方案) Entertainment AI solution drawing on 800TB of computer vision data and 2 million+ hours of speech data to enhance AI performance in gaming and entertainment.
- New Retail Solution (新零售解决方案) New retail data solution providing precise, efficient data for intelligent, personalized shopping experiences.
- Smart Healthcare Solution (智能医疗解决方案) Healthcare AI solution supporting multi-type data annotation through the proprietary platform and professional medical annotation teams for auxiliary diagnosis and health monitoring applications.
Quantifiable outcome
- Per capita annotation efficiency improved by 30%+
- +1 more outcomes
Companies that use Datatang
Customer profileNamed customers1 record
Segments7 records
Ideal customer profiles2 records
Datatang technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- No
- API docs
- API detail
Core technology
AI maturity
App detail
AI capability10 records
Feature6 records
Datatang partnerships and signals
Strategic signalPartnerships
Three partnerships are on record, tiered minor.
- openMPDminorOpenMPD is listed as a friendly link on Datatang's website footer, indicating some form of partnership or resource sharing relationship in the data ecosystem.
- 数加加 (Shujiajia)minorShujiajia is a related data annotation platform brand, listed as a friendly link. It appears to be an affiliated or sister platform for data annotation services.
- 帕依提提 (Payititi)minorPayititi is listed as a friendly link, suggesting a partnership or collaboration in the data services ecosystem.
Scale indicators4 records
Recent moves6 records
Expansion highlights5 records
Datatang competitors and assessment
Company assessmentDirect peers
- Labelbox: Labelbox offers a data labeling platform with managed annotation services across computer vision and NLP use cases. It competes directly with Datatang's Shujiajia annotation platform and annotation service offering.
- Appen: Appen provides human-annotated datasets for machine learning and AI across speech, text, image, and video modalities, serving global enterprise customers. It directly competes with Datatang in copyright dataset licensing and custom annotation services.
- Scale AI: Scale AI is a global leader in training data for AI, offering data annotation, evaluation, and RLHF services across computer vision, NLP, and autonomous driving. It is the most direct competitor to Datatang's full-stack data services model.
- Clickworker: Clickworker offers crowdsourced data collection and annotation services for AI training, including text, image, and audio labeling. It is a comparable peer in the custom data and crowdsourced annotation space.
- CloudFactory: CloudFactory provides managed data annotation and labeling workforces for AI/ML teams across computer vision and NLP. It is a direct competitor in the custom data annotation services segment Datatang serves.
- Defined.ai: Defined.ai operates an AI training data marketplace with pre-built datasets and custom data services across speech, vision, and NLP. It directly parallels Datatang's copyright dataset and customization offering.
Emerging players
- Snorkel AI: Snorkel AI uses programmatic labeling and weak supervision to reduce reliance on manual annotation, serving enterprise ML teams. It represents an adjacent, technology-driven threat to Datatang's manual annotation services.
- Turing: Turing provides AI training data services including RLHF and code labeling through a distributed expert workforce. It is an emerging player with partial overlap in LLM and code training data services adjacent to Datatang's offerings.
- Surge AI: Surge AI focuses on high-quality NLP data labeling and RLHF for foundation models, a fast-growing niche within AI training data. It represents emerging competition in Datatang's LLM dataset and annotation business.
Broad incumbents
- Amazon SageMaker Ground Truth: AWS SageMaker Ground Truth is a broad cloud platform offering managed data labeling services integrated with the AWS ecosystem. It competes as a broad incumbent in the same data annotation category Datatang serves.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat4 records
Key risks6 records
Key highlights6 records
Customer concentration
Datatang compliance and trust
Trust signalCompliance3 records
Datatang financial estimates
Financial estimateRevenue estimate
Valuation estimate
Datatang leadership team
Management profileNumber of profiles
Profiles1 record
Datatang funding detail
Funding detailFunding overview
Funding rounds1 record
Investors1 record
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Datatang M&A and investment
M&A and investmentM&A
Investments1 record
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Datatang
What does Datatang do?
Datatang provides copyright-cleared AI training datasets, customized data collection and annotation services, vertical-specific industry solutions, and its self-developed Shujiajia annotation platform for buyers developing AI and machine learning models. Offerings span large language model, computer vision, speech recognition/synthesis, OCR, and embodied intelligence data, with subject authorization that lets customers deploy the data immediately.
Is Datatang a public or private company?
Datatang is a private company. It is classified as founder individual operated bootstrapped and is currently operating.
When was Datatang founded?
Datatang was founded in 2010. It employs 501 to 1,000 people.
Where is Datatang based?
Datatang is headquartered in Beijing, China, in the Asia region.
How does Datatang make money?
Four revenue lines are on record. Copyright Dataset Sales are the primary driver. The others are data Customization Services, industry Solutions and data Annotation Platform Access.
Who are Datatang's main competitors?
Direct peers on record are Labelbox, Appen, Scale AI, Clickworker, CloudFactory and Defined.ai. Emerging players are Snorkel AI, Turing and Surge AI. Amazon SageMaker Ground Truth is listed as a broad incumbent.
Does Datatang have an API?
No public API is recorded for Datatang.
What industry is Datatang in?
Datatang's product category is AI Training Data Services. Its primary akta.pro industry code is HDAAALAB, Data Labeling & Annotation Services, with a secondary code of HDAAACAJ, Data Pipelines for GenAI (Curation, Filtering, Deduplication, Copyright). Its NAICS code is 541511 and its SIC code is 7370.