Diffbot
Diffbot is a private AI company that runs its own full-web crawl on custom hardware to build a structured Knowledge Graph of 10B+ entities, exposing it via Extract, Crawl, Natural Language, and Knowledge Graph APIs to 400+ enterprise and developer customers across finance, consumer, news, and risk verticals.
- Company typePrivate
- Founded2012
- HeadquartersMenlo Park, United States
- Headcount11–50
- GTM typeB2B
- OfferingSoftware
What Diffbot does
Diffbot Technologies Corp. is a private AI company founded in 2012 and headquartered at 333 Ravenswood Ave, Menlo Park, California. The company operates an independent full-web crawl on custom-assembled hardware in its own California datacenter, processing 1.2 billion public websites into a structured Knowledge Graph of 10B+ entities and 1T+ facts covering organizations (246M+), articles (1.6B+), retail products (3M+), discussions, and events (23k+). The pipeline applies computer vision for page-type classification and layout analysis, NLP for entity, relation, and sentiment extraction, and knowledge fusion to resolve entities across 20 languages. Underlying extraction technology is patented.
The platform is exposed as four API products — Extract (rule-less webpage structuring), Crawl (site-to-database bulk extraction), Natural Language (text-to-entities with claimed best-in-class accuracy versus Google, Microsoft, IBM, Amazon, OpenAI, and Stanford APIs), and Knowledge Graph (DQL search plus Enhance for record enrichment) — plus a sales-prospecting sub-brand called LeadGraph. Bundled Solutions target Market Intelligence, News Monitoring, Machine Learning training data, and Ecommerce use cases. Pricing is tiered: Free ($0/mo, 10K credits), Startup ($299/mo), Plus ($899/mo), and custom Enterprise, with usage-based credit overage billing; a Diffbot for Students program provides free Startup-tier access to academics.
Diffbot reports serving 400+ customers across finance (AlphaSense, Factset, FINRA, Diligent, Georgian, Valor Equity), consumer internet (Snapchat, Quora, Notion, Opera, Indeed, Klarna, Brex, AstraZeneca, Doximity), news and media intelligence (Meltwater, Cision, BusinessWire, NBC, BuzzFeed, SmartNews, SproutSocial, SemRush, IDC), and risk and compliance (ISS Governance, MerkleScience, Sigma Ratings, Orbital Insight, Riskwolf, Klue). Distribution combines a product-led self-serve funnel (free tier, marketplace add-ins for Google Sheets and Excel, Tableau and Power BI connectors, Zapier app, CData integration, public REST API with OpenAPI spec) with an enterprise field-sales motion via bespoke demo, [email protected], and 1-855-885-4800. Revenue is generated through recurring subscriptions, usage-based overages, data licensing of Search Subject records under CCPA/CPRA-compliant terms, and managed enterprise solutions. The company has raised $13M of disclosed venture funding across three rounds between 2012 and 2016; no subsequent funding activity is disclosed. CEO is Mike Tung; CTO is Arvid Sahlin.
Diffbot firmographics
Firmographics- Name
- Diffbot
- Legal name
- Diffbot Technologies Corp.
- Website
- https://diffbot.com
- Company type
- Private
- Founded year
- 2012
- Operating status
- Operating
- Headcount range
- 11–50 employees
- Short description
- Diffbot is a private AI company that runs its own full-web crawl on custom hardware to build a structured Knowledge Graph of 10B+ entities, exposing it via Extract, Crawl, Natural Language, and Knowledge Graph APIs to 400+ enterprise and developer customers across finance, consumer, news, and risk verticals.
- Ownership category
- akta.pro rank
Diffbot industry classification
Industry- Product category
- Web Data Infrastructure
- NAICS
- Web Search Portals and All Other Information Services (519290), Web Search Portals and All Other Information Services (51929), Web Search Portals, Libraries, Archives, and Other Information Services (5192), Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (518210), Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (5182)
- SIC
- Services-Prepackaged Software (7372), Services-Computer Processing & Data Preparation (7374), Services-Computer Programming Services (7371), Services-Computer Integrated Systems Design (7373)
- akta.pro primary industry
- Enterprise AI Data & Knowledge Platforms (Vector Databases, Knowledge Graphs) (HDAEANAH)
- akta.pro secondary industries
- Enterprise Search, Indexing & Content Discovery (HDAEAGAH), Search, Discovery & Ranking Personalization (HDAAAGAE), Search, Retrieval & Semantic Ranking (BM25/vector, hybrid) (HDAAADAB), Recommendation & Discovery Engines (content/product) (BPAMAAAG)
Keywords
Where Diffbot is headquartered
LocationHeadquarters
- HQ city
- Menlo Park
- HQ country
- United States
- HQ region
- North America
Offices2 records
Markets served
Diffbot business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Personnel, Technology or R&D, Infrastructure, Operations, Marketing or Sales, Others
Revenue model
- Subscription plans (Free / Startup / Plus / Enterprise): Tiered monthly subscriptions with included credit allotments (10K free; 250K Startup $299/mo; 1M Plus $899/mo; custom Enterprise), billed monthly with no long-term contract required.
- Usage-based credit overages: Pro-rata overage billing for credits consumed above the plan allotment at the plan's per-credit rate (e.g., $0.0009/credit on Plus, $0.001/credit on Startup).
- Freemium self-serve funnel: Permanent free plan with 10,000 credits/month, no credit card required, full API access and dashboard — used as a land motion to convert paid subscribers.
- Data licensing / sale of Search Subject records: Sale of personal information (name, employer, title, email, phone, social profile, education/employment history) from Search Subjects in the Knowledge Graph to Subscribers for B2B sales, marketing and recruiting use cases under license agreements.
- Custom managed solutions / bespoke enterprise builds: Enterprise tier offers bespoke plans and managed solutions where Diffbot builds custom data solutions for clients with no engineering resources.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Freemium | Monthly | Free tier: $0/mo, 10,000 credits, 5 calls/min, all products included, chat/email support. |
| Subscription | Monthly | Startup: $299/mo, 250,000 credits at $0.001/credit, 5 calls/sec. |
| Subscription | Monthly | Plus: $899/mo, 1,000,000 credits at $0.0009/credit, 25 calls/sec, 25 active crawls, 3 user licenses. |
| Other | Multi-year contract | Enterprise: Custom pricing, custom credit allotment and rate, 25+ calls/sec, 100+ active crawls, custom user licenses, premium SLA, managed solutions, shared Slack channel. |
| Usage-based | Pay-as-you-go | Usage-based credit overage billing on top of any paid plan. |
Go-to-market motion4 records
Distribution channels6 records
Marketing channels9 records
Diffbot product offering
Product offeringCore offering
Diffbot runs its own crawl of the entire public web on custom-assembled hardware in a California datacenter and uses AI (computer vision, NLP, and knowledge fusion) to transform unstructured web pages into a queryable Knowledge Graph of 10B+ entities and 1T+ facts. It sells access to this graph through four core API products — Extract (rule-less page structuring), Crawl (web-scale crawling), Natural Language (entity/relation/sentiment extraction), and Knowledge Graph Search/Enhance — plus a LeadGraph sub-brand for sales prospecting and four pre-built Solutions (Market Intelligence, News Monitoring, Machine Learning, Ecommerce), delivered via REST APIs, a self-serve dashboard, and add-ins for Excel, Google Sheets, Tableau, Power BI, and Zapier.
Product overview
Diffbot offers a unified Knowledge-as-a-Service platform built on a single product architecture: a core Knowledge Graph of the public web (10+ billion entities, 1 trillion facts) is populated and queried through four tightly coupled API products — Extract (rule-less webpage structuring), Crawl (web-scale crawling and bulk extraction), Natural Language (entity/relation/sentiment extraction from raw text), and Knowledge Graph (Search via DQL plus Enhance for record enrichment). The LeadGraph sub-brand extends the platform for sales prospecting, and four pre-built Solutions (Market Intelligence, News Monitoring, Machine Learning, Ecommerce) bundle the APIs into use-case workflows. All products share the same AI pipeline (computer vision, NLP, knowledge fusion) and are accessed via the same dashboard, REST APIs, and add-ins for Excel, Google Sheets, Tableau, Power BI, Zapier, and CData.
Differentiator
Problem solved
Functional benefit
Brands
- LeadGraph: Lead generation and prospecting data product offered by Diffbot, accessible at leadgraph.com and listed under the Diffbot products menu. Diffbot's California Privacy Statement also references 'Diffbot Dashboard or Leadgraph web app account login.'
Products and services
- Extract API Rule-less web data extraction API that automatically analyzes and structures articles, products, discussions, images, videos, and events into clean, normalized data fields using computer vision and visual layout analysis, with no per-site templates required. Built for developers and data teams who need structured data from any public webpage.
- Crawl API Web-scale crawler that turns any site into a structured database of products, articles, and discussions, supporting sitemaps, subdomains, recurring schedules, and bulk extraction. Built for data teams who need automated, continuous collection of structured data across entire domains.
- Natural Language API NLP API that infers entities, relationships, sentiment, and categories from raw text, with trainable customization for domain-specific entity and relationship sets. Supports multiple languages and is used by Volkswagen, Steelcase, JSTOR, UMass Amherst, and University of Alberta for research, document analysis, and content experiences.
- Knowledge Graph The world's largest automated Knowledge Graph of the public web, containing 10+ billion entities (organizations, articles, products, discussions, events, people, places, images, video, creative works) and 1+ trillion facts, searchable via Diffbot Query Language (DQL) and enrichable via the Enhance API. Forms the data backbone for Market Intelligence, News Monitoring, and LeadGraph use cases.
- LeadGraph LeadGraph is Diffbot's sub-brand product for sales prospecting and lead generation, powered by the Diffbot Knowledge Graph and accessible via leadgraph.com. Built for B2B sales teams that need qualified company and people data for outbound campaigns.
- Market Intelligence Solution Pre-built solution bundling Extract, Crawl, Natural Language, and Knowledge Graph to deliver 360-degree views of customers, suppliers, and competitors via DQL querying, fuzzy matching, and relationship discovery. Used by finance, consulting, and corporate-strategy teams for competitive monitoring and market mapping.
- News Monitoring Solution Pre-built solution for entity-aware news monitoring, sentiment tracking, and real-time alerting, powered by the Knowledge Graph news index and the Natural Language API. Built for PR, media intelligence, and brand-monitoring teams that need entity-disambiguated coverage across 20 languages.
- Machine Learning Solution Pre-built solution offering web-scale crawl data and Knowledge Graph exports for training ML models, including multi-lingual text, images, and structured data at billions-of-records scale. Built for ML practitioners, data scientists, and research organizations needing continuously updated training corpora.
- Ecommerce Solution Pre-built solution combining Extract, Crawl, and Knowledge Graph to monitor product listings, prices, inventory, reviews, and unauthorized sellers across the web. Built for brands, retailers, and ecommerce platforms needing price monitoring, MAP compliance, catalog enrichment, and competitive assortment analysis.
- Events Data Type (Knowledge Graph) Newly launched (2025) Events data type added to the Diffbot Knowledge Graph, with complete descriptions and normalized start/end timestamps; 23k+ events indexed at launch and extractable on demand via the Extract API.
- AI-powered Extract API update (2026) AI-powered update to Diffbot's Extract API launched in April 2026, generally available to subscribers and highlighted as a key industry development in the web scraping market.
Quantifiable outcome
- Knowledge Graph contains 10B+ entities and 1T+ facts (world's largest per company)
- +7 more outcomes
Companies that use Diffbot
Customer profileNamed customers35 records
Segments7 records
Ideal customer profiles6 records
Diffbot technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration6 records
AI capability9 records
Feature7 records
Diffbot partnerships and signals
Strategic signalPartnerships
18 partnerships are on record, tiered flagship and core.
- Google (Google Sheets / Google Workspace)flagshipDiffbot ships an official Google Sheets Add-On (Diffbot for Sheets) in the Google Workspace Marketplace that lets users query the Knowledge Graph and run Enhance from inside spreadsheets.
- Microsoft (Excel / Microsoft Graph)flagshipDiffbot publishes a Microsoft Excel Add-In that brings Knowledge Graph Search and Enhance into Excel workbooks. SSO is also offered via Microsoft Graph sign-in.
- Tableau (Salesforce)flagshipDiffbot Knowledge Graph is integrated with Tableau as a native data connector, letting users pull live standardized data on organizations and people into Tableau dashboards.
- ZapiercoreDiffbot Enhance is published as a Zapier app, enabling no-code workflow automation with thousands of other apps.
- CData
- Amazon Web Services (S3, EC2)flagshipDiffbot uses Amazon S3 for object storage and EC2 for on-demand processing capacity to support its on-demand web extraction services; processes data in the United States.
- StripeflagshipStripe is Diffbot's payment processor for paid subscriptions, self-certified under the Data Privacy Framework; PCI-DSS compliant.
- SalesforcecoreSalesforce CRM is used by Diffbot to unify prospect and customer data; data shared under a Data Protection Addendum.
- IntercomcoreIntercom powers Diffbot's in-app customer support messaging, product analytics and email communications.
- SlackcoreSlack is used as Diffbot's internal communication and collaboration platform; data transferred under the Data Privacy Framework.
- Plausible AnalyticscoreDiffbot self-hosts an open-source instance of Plausible Analytics for privacy-friendly visitor analytics; no data sent to third parties.
- AmplitudecoreAmplitude is used for cloud-based product analytics to help understand user behavior.
- MixMaxcoreMixMax provides Diffbot's outbound email platform, including templates, scheduling and engagement analytics for outreach to trial users and subscribers.
- Paddle (formerly Profitwell)corePaddle provides B2B subscription revenue automation, reporting and analytics for Diffbot; does not sell or share subscriber data.
- ReadMecoreReadMe powers Diffbot's documentation hub and developer platform at docs.diffbot.com.
- JSTORcoreAcademic publisher JSTOR worked with HBO and Diffbot to process rare primary interview transcripts for the Martin Luther King Jr. documentary 'King In The Wilderness', linking spoken word to structured knowledge.
- DianomicoreLargest native ad network in finance (350+ publishers) uses Diffbot NLP to monitor topics and analyze sentiment for brand-safe ad placements.
- LeadGraphflagshipLeadGraph is Diffbot's own sales-intelligence product (leadgraph.com) built on top of the Diffbot Knowledge Graph.
Scale indicators10 records
Recent moves6 records
Expansion highlights6 records
Diffbot competitors and assessment
Company assessmentDirect peers
- Zyte: Provider of web-scraping tools and services including Zyte API and Scrapy Cloud. Forbes (Nov 2025) lists Zyte as a comparable player in the same web-data infrastructure category as Diffbot.
- Apify: Cloud platform for web scraping and automation with a marketplace of 'actors' for data extraction. Forbes (Nov 2025) names Apify as a peer in the web-data infrastructure market.
- Common Crawl: Non-profit that maintains an open repository of web crawl data widely used as training input for LLMs. Diffbot's proprietary crawl infrastructure competes directly with Common Crawl's open dataset approach for ML training corpora.
- Bright Data: One of the largest web-data infrastructure platforms offering proxy networks, datasets, and scraping APIs. Forbes (Nov 2025) names Bright Data alongside Diffbot in the $209B web-data infrastructure market.
- Oxylabs: Web-scraping infrastructure provider offering proxies, scraping APIs, and ready-made datasets. Identified alongside Diffbot in the Forbes web-data infrastructure market landscape.
- ScrapingBee: Web-scraping API service that handles headless browsers and proxy rotation for developers. Competes with Diffbot's Extract API for developer-driven structured data extraction use cases.
- ScraperAPI: Proxy and web-scraping API platform targeting developers and data teams. Comparable to Diffbot's Extract and Crawl products in delivery model and developer customer base.
Broad incumbents
- AlphaSense: Enterprise market-intelligence platform that is both a Diffbot customer (using the Knowledge Graph for data feeds) and a broader incumbent building proprietary content sets for the same finance/market-intelligence buyer.
- Microsoft Bing: Microsoft's web search engine with its own index and entity knowledge. Diffbot's value proposition rests on operating an independent crawl and Knowledge Graph separate from Bing as well as Google.
- Google Knowledge Graph: Google's structured knowledge base powering Search features. Diffbot explicitly differentiates by running an independent crawl separate from Google, and Diffbot claims a news index 50x the size of Google's.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat6 records
Key risks7 records
Key highlights7 records
Customer concentration
Diffbot social profiles
Digital presenceDiffbot compliance and trust
Trust signalCompliance3 records
Diffbot financial estimates
Financial estimateRevenue estimate
Valuation estimate
Diffbot leadership team
Management profileNumber of profiles
Profiles7 records
Diffbot subsidiaries and ownership
Company hierarchySubsidiaries1 record
Diffbot funding detail
Funding detailFunding overview
Funding rounds3 records
Investors8 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Diffbot M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Diffbot
What does Diffbot do?
Diffbot runs its own crawl of the entire public web on custom-assembled hardware in a California datacenter and uses AI (computer vision, NLP, and knowledge fusion) to transform unstructured web pages into a queryable Knowledge Graph of 10B+ entities and 1T+ facts. It sells access to this graph through four core API products — Extract (rule-less page structuring), Crawl (web-scale crawling), Natural Language (entity/relation/sentiment extraction), and Knowledge Graph Search/Enhance — plus a LeadGraph sub-brand for sales prospecting and four pre-built Solutions (Market Intelligence, News Monitoring, Machine Learning, Ecommerce), delivered via REST APIs, a self-serve dashboard, and add-ins for Excel, Google Sheets, Tableau, Power BI, and Zapier.
Is Diffbot a public or private company?
Diffbot is a private company. It is classified as venture growth investor backed and is currently operating.
When was Diffbot founded?
Diffbot was founded in 2012. It employs 11 to 50 people.
Where is Diffbot based?
Diffbot is headquartered in Menlo Park, United States, in the North America region.
How does Diffbot make money?
Five revenue lines are on record. Subscription plans (Free / Startup / Plus / Enterprise) is the primary driver. The others are usage-based credit overages, freemium self-serve funnel, data licensing / sale of Search Subject records and custom managed solutions / bespoke enterprise builds.
Who are Diffbot's main competitors?
Direct peers on record are Zyte, Apify, Common Crawl, Bright Data, Oxylabs, ScrapingBee and ScraperAPI. Broad incumbents are AlphaSense, Microsoft Bing and Google Knowledge Graph.
Does Diffbot have an API?
Yes. Diffbot offers a public, REST-based developer API suite including the Extract API (Article, Product, Discussion, Video, Image, Custom), Natural Language API, Knowledge Graph Search/Enhance APIs, and Crawl/Bulk Extract. Developers can call APIs to autonomously extract structured entities, facts, relationships, and sentiment from any webpage or raw text; query and enrich the Knowledge Graph; and integrate via dashboards, REST, and add-ins. Authentication is via a unique API key issued per subscriber. Rate limits vary by plan (e.g., 5 calls/min on Free, 5 calls/sec on Startup, 25+ calls/sec on Plus/Enterprise). SDKs/code samples are provided in Shell, Node, Ruby, PHP, and Python, and an OpenAPI/Swagger specification is available. Developer documentation is at docs.diffbot.com.
What industry is Diffbot in?
Diffbot's product category is Web Data Infrastructure. Its primary akta.pro industry code is HDAEANAH, Enterprise AI Data & Knowledge Platforms (Vector Databases, Knowledge Graphs), with a secondary code of HDAEAGAH, Enterprise Search, Indexing & Content Discovery. Its NAICS code is 519290 and its SIC code is 7372.