Lexical Computing
Lexical Computing generates and sells large-scale language data assets — word frequency lists, n-gram databases, and enriched lexicons built from proprietary corpora totaling 1 trillion words across 100+ languages — alongside corpus query and dictionary editing tools, serving software developers, dictionary publishers, and AI/NLP customers.
- Company typePrivate
- Founded2003
- HeadquartersPortslade, United Kingdom
- Headcount11–50
- GTM typeB2B
- OfferingSoftware
What Lexical Computing does
Lexical Computing, founded in 2003 and headquartered in Portslade, United Kingdom (with an operating entity in the Czech Republic), builds and sells large-scale linguistic data assets and the tooling to use them. Its core technology is a proprietary pipeline that collects and cleans authentic web text to construct text corpora totaling 1 trillion words across 100+ languages, with the largest English corpus at 90 billion words. From these corpora the company generates enriched products — word frequency lists, n-gram (bigram, trigram and larger) databases, and word databases/lexicons containing millions of unique items per language with POS tags, lemmas, probabilities, synonyms, collocations, example sentences, and morphological information. The company positions these datasets as inputs for language modeling (LLMs), NLP applications, and corpus-based research.
The commercial portfolio spans three categories. Self-serve software tools — Sketch Engine (corpus query and management), Lexonomy (online dictionary editor), OneClick Terms (terminology extraction), and Lexicom (a course in lexicography) — operate on a subscription/recurring basis with free trials. Downloadable data products — word frequency lists, n-gram databases, and enriched word databases — are sold as one-time perpetual licenses, starting at EUR 250 for academic/research use and EUR 2,500 for commercial use, with custom quotations for large or specialized enterprise databases. The third layer is bespoke professional services covering terminology extraction, document classification, data mining, information retrieval, and custom corpus generation, sold via direct quotation.
The company sells horizontally to three primary buyer types: (i) software developers and AI/NLP practitioners needing reliable, large-scale language data; (ii) dictionary and language-teaching publishers needing lexical databases and frequency information; and (iii) enterprise teams needing language processing solutions for content management. Distribution is exclusively digital — a self-serve website with sample downloads, free trial accounts, custom-quotation enterprise sales, and presence on LinkedIn, Facebook, X, LinkedIn groups, and Academia.edu. The company operates globally with no geographic restrictions disclosed, is privately held with no institutional funding on record, and maintains a small team (11-50 employees).
Lexical Computing firmographics
Firmographics- Name
- Lexical Computing
- Legal name
- Lexical Computing
- Website
- https://lexicalcomputing.com
- Company type
- Private
- Founded year
- 2003
- Operating status
- Operating
- Headcount range
- 11–50 employees
- Short description
- Lexical Computing generates and sells large-scale language data assets — word frequency lists, n-gram databases, and enriched lexicons built from proprietary corpora totaling 1 trillion words across 100+ languages — alongside corpus query and dictionary editing tools, serving software developers, dictionary publishers, and AI/NLP customers.
- Ownership category
- akta.pro rank
Lexical Computing industry classification
Industry- Product category
- Language Data and Corpus Linguistics Software
- NAICS
- Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (5182), Web Search Portals, Libraries, Archives, and Other Information Services (519)
- SIC
- Services-Computer Processing & Data Preparation (7374), Services-Prepackaged Software (7372)
- akta.pro primary industry
- Search, Retrieval & Semantic Ranking (BM25/vector, hybrid) (HDAAADAB)
- akta.pro secondary industry
- Vocabulary, Flashcards, SRS & Memorization Platforms (EDAFAOAF)
Keywords
Where Lexical Computing is headquartered
LocationHeadquarters
- HQ city
- Portslade
- HQ country
- United Kingdom
- HQ region
- Europe
Markets served
Lexical Computing business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Personnel, Technology or R&D, Infrastructure, Marketing or Sales, Operations
Revenue model
- Word Database Sales: One-time purchase of word frequency lists, n-gram databases, and lexicons in multiple languages with pricing based on specifications, corpus size, and intended use
- Custom Data Services: Custom language data extraction, terminology extraction, document classification, and information retrieval solutions tailored to customer requirements
- Software Tools: Online tools including Sketch Engine corpus query system, Lexonomy dictionary editor, and OneClick Terms terminology extraction
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| One time/ perpetual license | Pay-as-you-go | Research/Academic Use - starting price |
| One time/ perpetual license | Pay-as-you-go | Commercial Use - standard price |
Go-to-market motion1 record
Distribution channels4 records
Marketing channels5 records
Lexical Computing product offering
Product offeringCore offering
Lexical Computing supplies large-scale word databases, lexicons, n-gram databases, and word frequency lists generated from text corpora containing approximately 1 trillion words across 100+ languages. It also provides software tools for corpus querying and management (Sketch Engine), online dictionary editing (Lexonomy), terminology extraction (OneClick Terms), and lexicography training (Lexicom), and offers custom language data and information retrieval services. Its customers are primarily software developers, dictionary and language teaching material publishers, and enterprises needing reliable language data.
Product overview
Lexical Computing provides a portfolio of language data products and tools built around a unified linguistic analysis platform. The core offerings include Sketch Engine (a corpus query and management system), Lexonomy (an online dictionary editor), OneClick Terms (terminology extraction), and Lexicom (educational courses). Supporting these are downloadable data products including word frequency lists, n-gram databases, and enriched word databases/lexicons generated from text corpora containing over 1 trillion words in 100+ languages. The company also provides solutions in full-text search, terminology extraction, document classification and categorization, data mining, and information retrieval.
Differentiator
Problem solved
Functional benefit
Brands
- Sketch Engine: Corpus query and management system for language data analysis
- Lexonomy
- OneClick Terms
- Lexicom
Products and services
- Sketch Engine Corpus query and management system that allows users to search and analyze text corpora for word frequency data, n-grams, collocations, and other linguistic information. Targeted at software developers, researchers, and language professionals.
- Lexonomy Online dictionary editor for creating and managing digital dictionaries and lexicons. Targeted at lexicographers, dictionary publishers, and language teaching content creators.
- OneClick Terms Terminology extraction tool that automatically identifies and extracts technical terms from text documents. Targeted at enterprise data teams, translators, and terminology managers.
- Lexicom A Course in Lexicography and Lexical Computing providing educational training in dictionary creation and language data processing. Targeted at students and professionals in lexicography and computational linguistics.
- Word Frequency Lists High-quality frequency word lists generated from text corpora, available for multiple languages including English, Spanish, French, Arabic, Russian, Portuguese, and Hindi. Targeted at software developers, NLP teams, lexicographers, and publishers.
- N-gram Databases Bigram, trigram, and larger n-gram databases generated from web corpora, available for multiple languages with millions of unique n-grams and options for enrichment with POS tags, lemmas, and probabilities. Targeted at software developers, NLP teams, and computational linguists.
- Word Databases and Lexicons Large word databases, lexical data, word lists, and lexicons enriched with linguistic data such as synonyms, collocations, example sentences, and morphological information. Targeted at dictionary publishers, language teaching material publishers, and software developers.
Quantifiable outcome
- Corpus sizes up to 90 billion words for single languages enabling rare word and n-gram frequency data
Companies that use Lexical Computing
Customer profileNamed customers3 records
Segments3 records
Ideal customer profiles3 records
Lexical Computing technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- No
- API docs
- API detail
Core technology
AI maturity
App detail
AI capability2 records
Feature5 records
Lexical Computing partnerships and signals
Strategic signalScale indicators11 records
Recent moves4 records
Expansion highlights4 records
Lexical Computing competitors and assessment
Company assessmentOthers
- Linguistic Data Consortium (LDC): LDC is the leading North American distributor of linguistic corpora, lexicons, and annotations to academic and government buyers; it is a distribution-side peer serving the same corpus linguistics and NLP research customers.
- Common Crawl Foundation: Common Crawl provides open, petabyte-scale web crawl data that many LLM teams use directly as a training corpus; it is the closest open-data alternative to Lexical Computing's cleaned text corpora and shapes market expectations on data pricing.
- ELRA (European Language Resources Association): ELRA / ELDA is the principal European distributor and consortium for language resources including corpora and lexica; it operates in the same buyer ecosystem as Lexical Computing and overlaps on European-language data licensing.
Broad incumbents
- Hugging Face: Hugging Face operates a broad hub for NLP models and datasets including pre-built corpora and language data; it overlaps with Lexical Computing in NLP data distribution and dictionary/lexicon-adjacent assets, but at vastly larger scale and broader scope.
- MemoQ / SDL Trados: MemoQ and SDL Trados are leading translation environment tools that include terminology management and translation memory features, overlapping with OneClick Terms and Lexonomy for dictionary and terminology workflows used by translators and language service providers.
Direct peers
- CQPweb: CQPweb, from Lancaster University, is a self-hosted corpus query system using the IMS Open Corpus Workbench; it is a direct technical peer for Sketch Engine in academic corpus linguistics deployments.
- AntConc: Laurence Anthony's AntConc is a free, widely used desktop corpus toolkit providing concordances, word lists, and n-gram analysis; it competes head-to-head with Sketch Engine for academic and corpus linguistics users and constrains pricing in that segment.
- LancsBox: LancsBox is Lancaster's modern successor desktop corpus toolkit with corpus building, annotation, and analysis features; it competes with Sketch Engine across the same corpus linguistics workflow.
- WordSmith Tools: Mike Scott's WordSmith Tools is a long-standing paid corpus analysis toolkit offering word frequency lists, concordances, and keyword analytics — a near-direct functional peer to Sketch Engine for the same lexicographer and corpus linguist buyer.
- Voyant Tools: Voyant Tools is a web-based, open-source text analysis environment for corpus reading and exploration; developed in the same digital humanities and corpus linguistics tradition, it overlaps directly with Sketch Engine's web-delivered analytical use cases.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat3 records
Key risks6 records
Key highlights7 records
Customer concentration
Lexical Computing social profiles
Digital presenceLexical Computing financial estimates
Financial estimateRevenue estimate
Valuation estimate
Lexical Computing leadership team
Management profileNumber of profiles
Lexical Computing funding detail
Funding detailFunding overview
Funding rounds
Investors
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Lexical Computing M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Lexical Computing
What does Lexical Computing do?
Lexical Computing supplies large-scale word databases, lexicons, n-gram databases, and word frequency lists generated from text corpora containing approximately 1 trillion words across 100+ languages. It also provides software tools for corpus querying and management (Sketch Engine), online dictionary editing (Lexonomy), terminology extraction (OneClick Terms), and lexicography training (Lexicom), and offers custom language data and information retrieval services. Its customers are primarily software developers, dictionary and language teaching material publishers, and enterprises needing reliable language data.
Is Lexical Computing a public or private company?
Lexical Computing is a private company. It is classified as unknown and is currently operating.
When was Lexical Computing founded?
Lexical Computing was founded in 2003. It employs 11 to 50 people.
Where is Lexical Computing based?
Lexical Computing is headquartered in Portslade, United Kingdom, in the Europe region.
How does Lexical Computing make money?
Three revenue lines are on record. Word Database Sales are the primary driver. The others are custom Data Services and software Tools.
Who are Lexical Computing's main competitors?
Others on record are Linguistic Data Consortium (LDC), Common Crawl Foundation and ELRA (European Language Resources Association). Broad incumbents are Hugging Face and MemoQ / SDL Trados. Direct peers are CQPweb, AntConc, LancsBox, WordSmith Tools and Voyant Tools.
Does Lexical Computing have an API?
No public API is recorded for Lexical Computing.
What industry is Lexical Computing in?
Lexical Computing's product category is Language Data and Corpus Linguistics Software. Its primary akta.pro industry code is HDAAADAB, Search, Retrieval & Semantic Ranking (BM25/vector, hybrid), with a secondary code of EDAFAOAF, Vocabulary, Flashcards, SRS & Memorization Platforms. Its NAICS code is 5182 and its SIC code is 7374.