Reworkd
- Company typePrivate
- Founded2023
- HeadquartersSan Francisco, United States
- Headcount1–10
- GTM typeB2B
- OfferingSoftware
Reworkd firmographics
Firmographics- Name
- Reworkd
- Legal name
- Reworkd AI, Inc.
- Website
- https://reworkd.ai
- Company type
- Private
- Founded year
- 2023
- Operating status
- Operating
- Headcount range
- 1–10 employees
- Ownership category
- akta.pro rank
Reworkd industry classification
Industry- Product category
- Web Data Extraction Platform
- NAICS
- Custom Computer Programming Services (541511), Document Preparation Services (561410)
- SIC
- Services-Prepackaged Software (7372), Services-Computer Programming, Data Processing, Etc. (7370)
- akta.pro primary industry
- LLMOps & Generative AI Platforms (Prompt/Agent Orchestration, RAG) (HDAEANAD)
- akta.pro secondary industry
- Process Mining & Task Mining Platforms (HDAEAHAD)
Keywords
Where Reworkd is headquartered
LocationHeadquarters
- HQ city
- San Francisco
- HQ country
- United States
- HQ region
- North America
Offices6 records
Markets served
Reworkd business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Technology or R&D, Personnel, Infrastructure, Marketing or Sales, Operations
Revenue model
- Subscription Plans: Three-tier subscription model: Hobby (free), Pro ($99/month), and Enterprise (custom pricing). Includes base features like concurrent browsers, data retention, API access, captcha solving, scheduled jobs, and fully managed infrastructure.
- Usage-Based Add-ons: Additional consumption charges for Standard Proxies ($0.125/GB), Premium Proxies ($8/GB), Compute ($0.10/hour), and Antibot Solving ($5/1,000 solved). These scale with customer usage beyond base plan allocations.
- Fully Managed Solution: Premium managed service including ongoing QA and maintenance of generated scraping code and resulting data. Includes dedicated Slack support and removes scraper maintenance burden from customer teams. Offered as part of Enterprise tier or add-on.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Freemium | Monthly | Free tier with basic features for individual developers to try the platform |
| Subscription | Monthly | Professional tier with expanded capacity and features for growing teams |
| Subscription | Multi-year contract | Enterprise tier with unlimited capacity and dedicated support |
Go-to-market motion3 records
Distribution channels4 records
Marketing channels7 records
Reworkd product offering
Product offeringCore offering
Reworkd operates an AI-powered web data extraction platform that automates the entire scraping pipeline—from scanning websites and generating code, to running extractors, validating results, and outputting structured data. The system uses LLM agents to dynamically generate and maintain Playwright scraping code based on customer-defined schemas, with self-healing capabilities that automatically repair failures when websites change.
Product overview
Reworkd is a unified AI-powered web data extraction platform offering end-to-end automation of web scraping pipelines. The core Reworkd platform uses LLM agents to generate and maintain Playwright scraping code based on customer-defined schemas. The product portfolio includes: the main Reworkd platform (core product), the Reworkd REST API for programmatic data access, the Harambe SDK for developers, and AgentGPT (legacy general-purpose agent brand). Add-on modules include Templates (scraper reuse), Deduplication (change tracking), Scheduling (automated re-runs), API Exports and Bulk Exports (data retrieval), and File Downloads (automatic file handling). A partnership with NewsCatcher extends access to 90,000+ news sources. The platform supports multiple data types including text, images, documents, and structured data.
Differentiator
Problem solved
Functional benefit
Brands
- Harambe: A custom web scraping SDK built by Reworkd for saving data, validating schemas, enqueuing URLs, de-duplicating data, and handling web scraping problems like pagination, PDFs, and downloads.
- AgentGPT
Products and services
- Reworkd Platform AI-powered end-to-end web data extraction platform that automates the entire scraping pipeline. Uses LLM agents to generate and maintain Playwright scraping code based on customer-defined schemas. Available in Hobby (free), Pro ($99/month), and Enterprise (custom) tiers for businesses that need to extract web data at scale.
- Reworkd API Public REST API for programmatically exporting scraped data, managing cookies and local storage, and retrieving review status. Supports pagination, incremental sync via created_after filtering, and country/region filtering. Available on all paid tiers for technical teams integrating web data into their pipelines.
- Harambe SDK Custom Python SDK automatically generated as part of the scraping code generation workflow. Provides methods for saving data with schema validation, enqueuing URLs, handling pagination, capturing downloads, HTML/PDF capture, and logging. Built for developers extending Reworkd's extraction capabilities.
Quantifiable outcome
- Extract hundreds of thousands of regulation PDFs monthly, saving hundreds of engineering hours
- +7 more outcomes
Companies that use Reworkd
Customer profileNamed customers3 records
Segments5 records
Ideal customer profiles3 records
Reworkd technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
AI capability8 records
Feature13 records
Reworkd partnerships and signals
Strategic signalPartnerships
Two partnerships are on record, tiered core.
- NewsCatchercoreReworkd partners with NewsCatcher to provide customers with instant access to news data from 90,000+ sources. This partnership enhances Reworkd's data offering beyond standard web scraping by integrating NewsCatcher's near-real-time news enrichment, NLP techniques (sentiment analysis, entity detection), and precise tagging. The integration helps customers unlock value faster for AI model training and business intelligence.
- ConfluentcoreReworkd built an AI-powered real-time web scraping system using Confluent's Kafka-based data streaming platform. The integration enables near-real-time data processing, fault tolerance, and scalable data pipelines for downstream consumers. Built with OpenAI's GPT-4 via Azure for automated code generation and validation.
Scale indicators11 records
Recent moves6 records
Expansion highlights5 records
Reworkd competitors and assessment
Company assessmentDirect peers
- Apify: Apify is a web scraping and automation platform offering a marketplace of pre-built scrapers ("Actors") plus cloud infrastructure for running extraction jobs. Directly competes with Reworkd on the developer-facing, self-serve web data extraction use case.
- Zyte (formerly Scrapinghub): Zyte provides enterprise web data extraction services and tools, including the Crawlera proxy network and Zyte API for managed scraping. Competes with Reworkd in the enterprise tier serving customers needing large-scale, managed extraction across thousands of sites.
- Bright Data: Bright Data is a leading web data platform combining proxy networks, scraping APIs, and datasets. Its Web Unlocker and Scraping Browser overlap directly with Reworkd's proxy/antibot and browser automation capabilities.
- Diffbot: Diffbot offers AI-powered web data extraction via its Knowledge Graph and product/article APIs, competing with Reworkd on structured data extraction at scale without manual scraper maintenance.
- Octoparse: Octoparse is a no-code web scraping tool aimed at business users and SMBs, competing with Reworkd's Hobby and Pro tiers in the visual, no-code extraction market.
Emerging players
- Browserbase: Browserbase provides cloud-hosted headless browser infrastructure for AI agents and web automation, overlapping with Reworkd's underlying browser execution layer. Both target developers building agentic scraping workflows.
- ScrapeGraphAI: ScrapeGraphAI uses LLMs to convert natural language prompts into structured scraping pipelines, directly mirroring Reworkd's LLM-agent approach and competing for AI-native developers.
- Skyvern: Skyvern applies LLM-driven browser agents to automate web workflows beyond pure extraction, including form filling and navigation. Overlaps with Reworkd in agentic browser automation for enterprise use cases.
- MultiOn: MultiOn builds AI agents that autonomously operate web browsers to complete tasks, intersecting with Reworkd in the broader category of agentic browser automation. Both target the AI-agent market with overlapping underlying technology.
Broad incumbents
- Import.io: Import.io is an established enterprise data extraction platform with mature offerings for web data integration, datasets, and APIs. Competes with Reworkd's Enterprise tier for large organizations needing managed extraction at scale.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat5 records
Key risks6 records
Key highlights7 records
Customer concentration
Reworkd social profiles
Digital presenceReworkd financial estimates
Financial estimateRevenue estimate
Valuation estimate
Reworkd leadership team
Management profileNumber of profiles
Profiles3 records
Reworkd funding detail
Funding detailFunding overview
Funding rounds2 records
Investors7 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Reworkd M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Reworkd
What does Reworkd do?
Reworkd operates an AI-powered web data extraction platform that automates the entire scraping pipeline—from scanning websites and generating code, to running extractors, validating results, and outputting structured data. The system uses LLM agents to dynamically generate and maintain Playwright scraping code based on customer-defined schemas, with self-healing capabilities that automatically repair failures when websites change.
Is Reworkd a public or private company?
Reworkd is a private company. It is classified as venture growth investor backed and is currently operating.
When was Reworkd founded?
Reworkd was founded in 2023. It employs 1 to 10 people.
Where is Reworkd based?
Reworkd is headquartered in San Francisco, United States, in the North America region.
How does Reworkd make money?
Three revenue lines are on record. Subscription Plans are the primary driver. The others are usage-Based Add-ons and fully Managed Solution.
Who are Reworkd's main competitors?
Direct peers on record are Apify, Zyte (formerly Scrapinghub), Bright Data, Diffbot and Octoparse. Emerging players are Browserbase, ScrapeGraphAI, Skyvern and MultiOn. Import.io is listed as a broad incumbent.
Does Reworkd have an API?
Yes. Reworkd provides a public REST API accessible at https://api.reworkd.dev/v1 for exporting scraped data. The API supports Bearer token authentication. Key endpoints include: GET /v1/outputs/{group_id} for fetching scraping group outputs, GET/POST /v1/cookies for cookie management, GET/POST /v1/local-storage for local storage operations, and GET /v1/reviews/{group_id} for review status. The API supports pagination with cursor, filtering by created_after date for incremental syncs, and country/region filtering. Developer documentation is at docs.reworkd.ai/api-reference/public/get-outputs-for-a-scraping-group.
What industry is Reworkd in?
Reworkd's product category is Web Data Extraction Platform. Its primary akta.pro industry code is HDAEANAD, LLMOps & Generative AI Platforms (Prompt/Agent Orchestration, RAG), with a secondary code of HDAEAHAD, Process Mining & Task Mining Platforms. Its NAICS code is 541511 and its SIC code is 7372.