Scikit-Learn
Scikit-learn is a free, BSD-licensed open-source Python machine learning library offering classification, regression, clustering, dimensionality reduction, and model selection tools for data scientists, researchers, and enterprise teams worldwide, sustained by institutional sponsorship and community contributions.
- Company typePrivate
- Founded2007
- HeadquartersParis, France
- Headcount11–50
- GTM typeB2B
- OfferingSoftware
What Scikit-Learn does
Scikit-learn is an open-source machine learning library for Python, initiated in 2007 and headquartered in Paris, France, that provides supervised and unsupervised learning algorithms (classification, regression, clustering, dimensionality reduction), model selection (grid search, cross-validation), and preprocessing utilities. It is built on the core scientific Python stack (NumPy, SciPy, matplotlib) under a permissive BSD license that permits unrestricted commercial use, and is distributed via PyPI, conda-forge, GitHub, and source downloads to a global base of data scientists, researchers, and enterprise engineering teams.
The project is operated as a volunteer-driven community initiative without a formal corporate entity, organized under the Inria Foundation consortium and sponsored by a diverse group of institutions including the Chan Zuckerberg Initiative, Wellcome Trust, NVIDIA, Quansight Labs, NASA, Chanel, BNP Paribas Group, and Michelin. Governance, infrastructure, and core contributor coordination are handled through NumFOCUS-affiliated donation infrastructure, a community mailing list, Discord, GitHub Discussions, and recurring in-person community sprints held across multiple continents. Enterprise-grade support and services are offered through Probabl, an affiliated commercial company, which is the only sanctioned monetization channel tied to the open-source codebase.
There is no direct software revenue stream associated with scikit-learn proper; sustainability is maintained through institutional grants, philanthropic donations, and volunteer contributions. In 2024 the project received the Chan Zuckerberg Initiative's Essential Open Source Software for Science (EOSS6) award, validating its scientific impact, and it continues to ship a regular cadence of major and minor releases (versions 1.7.x and 1.8.0 in 2025, 1.9.0 in June 2026, with 1.10 in development).
Scikit-Learn firmographics
Firmographics- Name
- Scikit-Learn
- Legal name
- scikit-learn
- Website
- https://scikit-learn.org
- Company type
- Private
- Founded year
- 2007
- Operating status
- Operating
- Headcount range
- 11–50 employees
- Short description
- Scikit-learn is a free, BSD-licensed open-source Python machine learning library offering classification, regression, clustering, dimensionality reduction, and model selection tools for data scientists, researchers, and enterprise teams worldwide, sustained by institutional sponsorship and community contributions.
- Ownership category
- akta.pro rank
Scikit-Learn industry classification
Industry- Product category
- Machine Learning Software
- SIC
- Services-Prepackaged Software (7372)
- akta.pro primary industry
- Education & Learning Personalization (HDAAAGAL)
Keywords
Where Scikit-Learn is headquartered
LocationHeadquarters
- HQ city
- Paris
- HQ country
- France
- HQ region
- Europe
Markets served
Scikit-Learn business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Personnel, Technology or R&D, Operations
Revenue model
- Open Source Distribution (No Direct Revenue): Scikit-learn is freely distributed as open source software under the BSD license. The project generates no direct revenue from software licensing. Sustainability is maintained through institutional funding, donations, and volunteer contributions from the global developer community.
Go-to-market motion1 record
Distribution channels5 records
Marketing channels10 records
Scikit-Learn product offering
Product offeringCore offering
Scikit-learn is an open-source Python machine learning library built on NumPy, SciPy, and matplotlib. It provides simple and efficient tools for predictive data analysis, including supervised learning (classification, regression), unsupervised learning (clustering, dimensionality reduction), model selection, and preprocessing. The library is freely distributed under the BSD license for both academic and commercial use.
Product overview
Scikit-learn is a single unified open-source Python library for machine learning. The core product encompasses all its capabilities as one integrated package rather than a platform-with-modules architecture. The library provides tools organized into modules: sklearn.cluster (clustering algorithms like K-Means, HDBSCAN, hierarchical clustering), sklearn.ensemble (gradient boosting, random forests, AdaBoost), sklearn.linear_model (logistic/linear regression, Lasso, Ridge), sklearn.svm (support vector machines), sklearn.decomposition (PCA, NMF, ICA, dictionary learning), sklearn.model_selection (GridSearchCV, cross-validation), sklearn.preprocessing (feature scaling, encoding, discretization), sklearn.neighbors (nearest neighbors algorithms), sklearn.naive_bayes, sklearn.gaussian_process, sklearn.neural_network (MLP), sklearn.metrics, sklearn.impute, sklearn.feature_extraction, sklearn.feature_selection, sklearn.manifold (t-SNE, Isomap), sklearn.mixture (Gaussian Mixture Models), sklearn.covariance, sklearn.tree, sklearn.dummy, sklearn.kernel_approximation, sklearn.kernel_ridge, sklearn.isotonic, sklearn.semi_supervised, sklearn.multiclass, sklearn.multioutput, sklearn.calibration, sklearn.inspection, sklearn.frozen, sklearn.compose (ColumnTransformer, Pipeline), sklearn.cross_decomposition, sklearn.random_projection, sklearn.datasets, sklearn.callback, sklearn.utils, sklearn.base, sklearn.exceptions, sklearn.experimental, sklearn.config_context.
Differentiator
Problem solved
Functional benefit
Products and services
- scikit-learn
Quantifiable outcome
- BSD license enables unrestricted commercial use
- +1 more outcomes
Companies that use Scikit-Learn
Customer profileNamed customers6 records
Segments3 records
Ideal customer profiles3 records
Scikit-Learn technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- No
- API docs
- API detail
Core technology
AI maturity
App detail
AI capability12 records
Feature8 records
Scikit-Learn partnerships and signals
Strategic signalPartnerships
Two partnerships are on record, tiered minor and core.
- Hugging FaceminorMachine learning platform joining forces with scikit-learn community. Collaboration announced in October 2022 for ecosystem integration and joint community efforts.
- ProbablcoreEnterprise-grade solutions and services provider for scikit-learn. Offers commercial support, training, and services around the library. Affiliated with the core scikit-learn ecosystem.
Scale indicators4 records
Recent moves5 records
Expansion highlights5 records
Scikit-Learn competitors and assessment
Company assessmentDirect peers
- PyTorch: PyTorch is the dominant open-source deep learning framework from Meta. It competes with scikit-learn in the broader Python ML library category, particularly for predictive modeling and data analysis workflows, though it focuses on neural networks rather than classical ML.
- Hugging Face: Hugging Face is the leading community platform for modern ML models, transformers, and datasets. It has an active community collaboration with scikit-learn (announced October 2022) and serves overlapping developer audiences building ML applications in Python.
- XGBoost: XGBoost is a popular open-source gradient boosting library. It is often used alongside scikit-learn (and now integrated into sklearn's HistGradientBoosting), targeting the same data science practitioners for predictive modeling competitions and enterprise tabular ML.
- LightGBM: LightGBM is Microsoft's gradient boosting framework optimized for performance and scale. It directly competes with scikit-learn's gradient boosting implementations for tabular ML use cases where speed and scale matter.
- Keras: Keras is a high-level deep learning API that runs on top of TensorFlow, PyTorch, or JAX. While focused on neural networks, it competes for the same Python ML practitioner mindshare that scikit-learn has historically dominated.
Broad incumbents
- TensorFlow: TensorFlow is Google's end-to-end open-source ML platform. While broader in scope than scikit-learn, it overlaps significantly in classification, regression, and production ML workflows, and represents the incumbent enterprise choice in many large organizations.
Emerging players
- Apache MXNet: Apache MXNet is an Apache Software Foundation deep learning framework. It overlaps with scikit-learn as part of the broader Python ML ecosystem, though with different primary use cases and adoption levels.
- Probabl: Probabl is a commercial spin-out offering enterprise-grade solutions and services around scikit-learn. It is the closest commercial entity affiliated with the scikit-learn ecosystem and provides a paid support and services channel.
Others
- NumPy: NumPy is the foundational numerical computing library that scikit-learn is built upon. It is an ecosystem enabler rather than a direct competitor, but the two are tightly coupled in nearly all scikit-learn workflows.
- Pandas: Pandas is the standard Python data manipulation and analysis library. It is a complementary tool that scikit-learn users almost universally pair with for data preprocessing and feature engineering workflows.
Market position
Strengths5 records
Weaknesses4 records
Competitive moat6 records
Key risks5 records
Key highlights7 records
Customer concentration
Scikit-Learn social profiles
Digital presenceScikit-Learn financial estimates
Financial estimateRevenue estimate
Valuation estimate
Scikit-Learn leadership team
Management profileNumber of profiles
Scikit-Learn funding detail
Funding detailFunding overview
Funding rounds
Investors
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Scikit-Learn M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Scikit-Learn
What does Scikit-Learn do?
Scikit-learn is an open-source Python machine learning library built on NumPy, SciPy, and matplotlib. It provides simple and efficient tools for predictive data analysis, including supervised learning (classification, regression), unsupervised learning (clustering, dimensionality reduction), model selection, and preprocessing. The library is freely distributed under the BSD license for both academic and commercial use.
Is Scikit-Learn a public or private company?
Scikit-Learn is a private company. It is classified as unknown and is currently operating.
When was Scikit-Learn founded?
Scikit-Learn was founded in 2007. It employs 11 to 50 people.
Where is Scikit-Learn based?
Scikit-Learn is headquartered in Paris, France, in the Europe region.
How does Scikit-Learn make money?
One revenue line is on record: open Source Distribution (No Direct Revenue).
Who are Scikit-Learn's main competitors?
Direct peers on record are PyTorch, Hugging Face, XGBoost, LightGBM and Keras. TensorFlow is listed as a broad incumbent. Emerging players are Apache MXNet and Probabl. Others are NumPy and Pandas.
Does Scikit-Learn have an API?
No public API is recorded for Scikit-Learn.
What industry is Scikit-Learn in?
Scikit-Learn's product category is Machine Learning Software. Its primary akta.pro industry code is HDAAAGAL, Education & Learning Personalization. Its SIC code is 7372.