spaCy
spaCy is an open-source Python NLP library published by Berlin-based Explosion since 2015, serving developers and ML engineers with production-grade tokenization, NER, parsing, and transformer-based pipelines across 75+ languages. Explosion monetizes via its Prodigy annotation tool, custom pipeline services, and the ellf.ai agentic NLP beta.
- Company typePrivate
- Founded2015
- HeadquartersBerlin, Germany
- Headcount—
- GTM typeB2B
- OfferingSoftware
What spaCy does
spaCy is an open-source Python library for industrial-strength natural language processing, published since 2015 by Explosion, a private Berlin-based company founded by Ines Montani. The library provides tokenization, part-of-speech tagging, dependency parsing, lemmatization, named entity recognition, entity linking, text classification, and rule-based matching across 75+ languages, with 84 pretrained pipelines for 25 languages. Its architecture is a configurable pipeline built on Cython for performance (benchmarks report 10,014 words/second on CPU and 14,954 on GPU with the large English pipeline), with support for custom models in PyTorch and TensorFlow, GPU acceleration via CuPy, and transformer integration through curated BERT/RoBERTa/GPT-2 pipelines.
The product portfolio extends beyond the core library: spacy-llm integrates OpenAI and Hugging Face large language models into spaCy pipelines via modular prompting and structured output parsing; spacy-transformers and spacy-curated-transformers provide state-of-the-art accuracy paths; Prodigy is Explosion's commercial annotation tool tightly coupled to spaCy training; spaCy Projects orchestrates end-to-end ML workflows with DVC, Weights & Biases, Streamlit, FastAPI, and Hugging Face Hub integrations; and the spaCy Universe catalogues 200+ community plugins, extensions, and educational resources.
Explosion does not monetize the spaCy library itself, which is free and open-source under a community-led, API-first go-to-market motion distributed via pip and conda-forge. Revenue derives from Prodigy (a recurring-license commercial annotation tool), fixed-fee Custom Solutions where Explosion's core developers build bespoke pipelines for enterprise clients, and emerging products such as the ellf.ai agentic NLP beta. Customer segments are horizontal and persona-based, spanning individual developers, data scientists, ML engineers, and production NLP teams globally.
spaCy firmographics
Firmographics- Name
- spaCy
- Legal name
- Explosion
- Website
- https://spacy.io
- Company type
- Private
- Founded year
- 2015
- Operating status
- Operating
- Short description
- spaCy is an open-source Python NLP library published by Berlin-based Explosion since 2015, serving developers and ML engineers with production-grade tokenization, NER, parsing, and transformer-based pipelines across 75+ languages. Explosion monetizes via its Prodigy annotation tool, custom pipeline services, and the ellf.ai agentic NLP beta.
- Ownership category
- akta.pro rank
Where spaCy is headquartered
LocationHeadquarters
- HQ city
- Berlin
- HQ country
- Germany
- HQ region
- Europe
Offices1 record
Markets served
spaCy business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Technology or R&D, Personnel, Infrastructure, Marketing or Sales, Operations
Revenue model
- Open-source library (free): spaCy is a free, open-source library distributed at no cost. The library itself does not generate direct revenue. The company behind spaCy (Explosion) monetizes through complementary commercial products and services.
- Prodigy (Commercial Annotation Tool): Explosion sells Prodigy, a radically efficient machine teaching and annotation tool. While separate from spaCy, it is tightly integrated and promoted as the annotation solution for training spaCy models.
- Custom Solutions Services: Explosion offers custom spaCy pipeline development services, where their core developers create tailor-made NLP pipelines for specific business needs. This is a professional services engagement.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Freemium | Annual | Free open-source library |
Go-to-market motion2 records
Distribution channels6 records
Marketing channels9 records
spaCy product offering
Product offeringCore offering
spaCy is an industrial-strength, open-source Python library for production-grade natural language processing, offering tokenization, POS tagging, dependency parsing, named entity recognition, entity linking, text classification, and lemmatization across 75+ languages. The publishing company (Explosion) monetizes through the commercial Prodigy annotation tool and custom NLP pipeline development services delivered by spaCy's core developers.
Product overview
spaCy is an open-source NLP library for Python built by Explosion, designed for production use. The core product is spaCy itself (the NLP library), which provides tokenization, POS tagging, dependency parsing, lemmatization, NER, entity linking, text classification, and more for 75+ languages. The product portfolio includes key add-on modules: spacy-llm integrates LLMs (GPT-4, GPT-3.5, Hugging Face models) into spaCy pipelines; spacy-transformers and spacy-curated-transformers provide transformer-based pipelines (BERT, RoBERTa); Prodigy is Explosion's own annotation tool for creating training data; spaCy Projects manages end-to-end workflows; spacy-streamlit provides Streamlit visualizations; and spacy-huggingface-hub enables model sharing on Hugging Face. The spaCy Universe catalogs 150+ community plugins and extensions. Custom Solutions (by spaCy's core developers) delivers tailored pipelines. An interactive online course and VS Code extension support developer onboarding.
Differentiator
Problem solved
Functional benefit
Brands
- spacy-llm: Package that integrates Large Language Models (LLMs) into spaCy pipelines
- Prodigy
- Thinc
Products and services
- spaCy Industrial-strength open-source NLP library for Python, providing tokenization, POS tagging, dependency parsing, lemmatization, NER, entity linking, text classification, and rule-based matching for 75+ languages, with 84 trained pipelines available for 25 languages. Designed for developers and data scientists building production NLP systems.
- Prodigy Commercial annotation tool developed by Explosion for creating training data for machine learning models, featuring active learning workflows for entity recognition, intent detection, and image classification. Integrates natively with spaCy for training and updating spaCy models.
- spaCy Custom Solutions Professional service by spaCy's core developers delivering custom, production-ready NLP pipelines tailored to specific business problems. Deliverables include a complete spaCy project folder with full code, data, tests, and documentation, quoted at a fixed fee up-front.
- spacy-llm
- spaCy Projects
- spaCy Universe
- spaCy Online Course
Quantifiable outcome
- 10,014 words processed per second on CPU with en_core_web_lg pipeline
- +3 more outcomes
Companies that use spaCy
Customer profileSegments3 records
Ideal customer profiles3 records
spaCy technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration26 records
AI capability9 records
Feature9 records
spaCy partnerships and signals
Strategic signalPartnerships
Five partnerships are on record, tiered core.
- ProdigycoreProdigy is an annotation tool developed by Explosion (the company behind spaCy). It is tightly integrated with spaCy and promoted as the recommended tool for creating training data to train or update spaCy models. Features active learning workflows for entity recognition, intent detection, and image classification. Prodigy is a commercial product that complements spaCy's open-source offering.
- OpenAIcorespacy-llm integrates OpenAI's GPT models (GPT-3.5, GPT-4, and variants) into spaCy pipelines. Users can configure OpenAI API access for LLM-powered NLP tasks within the spaCy framework. Requires API key configuration.
- Hugging FacecorespaCy integrates with Hugging Face Hub for model sharing. The spacy-huggingface-hub package enables uploading spaCy pipelines to Hugging Face. Additionally, spacy-llm supports open-source models hosted on Hugging Face (e.g., Dolly).
- PyTorch / TensorFlowcorespaCy supports custom models in PyTorch, TensorFlow, and other frameworks. The transformer-based pipelines use these frameworks for the underlying ML models.
- CuPycoreCuPy provides GPU acceleration support for spaCy. spaCy uses CuPy for CUDA-compatible GPU processing, enabling high-throughput text processing on GPUs.
Scale indicators7 records
Recent moves6 records
Expansion highlights6 records
spaCy competitors and assessment
Company assessmentDirect peers
- AllenNLP: AllenNLP is an open-source NLP research library built on PyTorch. It overlaps with spaCy's customizable components, model training, and pipeline extensibility for research and production NLP workloads.
- Flair: Flair is an open-source NLP framework built on PyTorch, providing state-of-the-art NER, POS tagging, and text classification. It is directly benchmarked against spaCy on NER tasks and serves similar developer and research personas.
- Hugging Face: Hugging Face provides open-source NLP/transformer libraries and a model hub that overlaps directly with spaCy's pretrained pipelines and Hugging Face Hub integration. Both target developers building NLP applications with state-of-the-art transformer models.
- NLTK (Natural Language Toolkit): NLTK is a foundational open-source Python NLP library used widely in academia. It overlaps with spaCy as a general-purpose NLP toolkit, though it is more research-oriented and slower for production use.
- Stanford Stanza (Stanford NLP Group): Stanza is a Python NLP library from Stanford offering tokenization, POS tagging, NER, and dependency parsing across 70+ languages. It is benchmarked head-to-head with spaCy and bridges research-grade models into Python pipelines.
Emerging players
- Rasa: Rasa is an open-source conversational AI platform that uses spaCy under the hood for NER and intent classification. It represents an emerging player building specialized applications on top of spaCy's NLP primitives.
- John Snow Labs (spark-nlp): John Snow Labs provides the spark-nlp library and enterprise NLP solutions for healthcare and regulated industries. It is an emerging commercial player with overlapping production NLP positioning and pretrained model coverage.
- LangChain: LangChain is a framework for orchestrating LLM applications and integrates with spaCy via spacy-llm. It is an emerging player in the adjacent LLM orchestration space that competes for developer mindshare in applied NLP pipelines.
Broad incumbents
- Google Cloud Natural Language API: Google Cloud Natural Language API is a managed cloud NLP service offering entity analysis, sentiment analysis, and classification. It is a broad incumbent that competes with spaCy at the enterprise procurement layer for non-developer buyers.
- Amazon Comprehend: Amazon Comprehend is AWS's managed NLP service for entity extraction, sentiment, and topic modeling. It targets enterprise buyers who prefer managed services over assembling open-source NLP stacks like spaCy.
Market position
Strengths4 records
Weaknesses4 records
Competitive moat5 records
Key risks6 records
Key highlights7 records
Customer concentration
spaCy social profiles
Digital presencespaCy financial estimates
Financial estimateRevenue estimate
Valuation estimate
spaCy leadership team
Management profileNumber of profiles
Profiles1 record
spaCy funding detail
Funding detailFunding overview
Funding rounds
Investors
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
spaCy M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about spaCy
What does spaCy do?
spaCy is an industrial-strength, open-source Python library for production-grade natural language processing, offering tokenization, POS tagging, dependency parsing, named entity recognition, entity linking, text classification, and lemmatization across 75+ languages. The publishing company (Explosion) monetizes through the commercial Prodigy annotation tool and custom NLP pipeline development services delivered by spaCy's core developers.
Is spaCy a public or private company?
spaCy is a private company. It is classified as founder individual operated bootstrapped and is currently operating.
When was spaCy founded?
spaCy was founded in 2015.
Where is spaCy based?
spaCy is headquartered in Berlin, Germany, in the Europe region.
How does spaCy make money?
Three revenue lines are on record. Open-source library (free) is the primary driver. The others are prodigy (Commercial Annotation Tool) and custom Solutions Services.
Who are spaCy's main competitors?
Direct peers on record are AllenNLP, Flair, Hugging Face, NLTK (Natural Language Toolkit) and Stanford Stanza (Stanford NLP Group). Emerging players are Rasa, John Snow Labs (spark-nlp) and LangChain. Broad incumbents are Google Cloud Natural Language API and Amazon Comprehend.
Does spaCy have an API?
Yes. spaCy exposes a Python API (not a traditional REST API) for programmatic access. It provides a CLI (spacy CLI commands) and Python library for training, evaluation, and pipeline management. For REST API access, third-party packages are used: spacy-api-docker (Docker-based REST API wrapper), FastAPI integration (spaCy+FastAPI template), spacy-nlp (Node.js via Socket.IO), spacy-js (JavaScript API), spacy-graphql (GraphQL interface). spacy-llm integrates LLMs into spaCy pipelines supporting OpenAI API (GPT-4, GPT-3.5) and Hugging Face open-source models. Developer documentation is at spacy.io/api.