Developer docs
API playgroundTry for free, no card

Search company profiles

Lexical Computing

Full company profile

uuid000lgbf

Namestring
Lexical Computing
Legal namestring
Lexical Computing
Company typeenum
Private
Founded yearint
2003
Descriptiontext

Lexical Computing, founded in 2003 and headquartered in Portslade, United Kingdom (with an operating entity in the Czech Republic), builds and sells large-scale linguistic data assets and the tooling to use them. Its core technology is a proprietary pipeline that collects and cleans authentic web text to construct text corpora totaling 1 trillion words across 100+ languages, with the largest English corpus at 90 billion words. From these corpora the company generates enriched products — word frequency lists, n-gram (bigram, trigram and larger) databases, and word databases/lexicons containing millions of unique items per language with POS tags, lemmas, probabilities, synonyms, collocations, example sentences, and morphological information. The company positions these datasets as inputs for language modeling (LLMs), NLP applications, and corpus-based research.

The commercial portfolio spans three categories. Self-serve software tools — Sketch Engine (corpus query and management), Lexonomy (online dictionary editor), OneClick Terms (terminology extraction), and Lexicom (a course in lexicography) — operate on a subscription/recurring basis with free trials. Downloadable data products — word frequency lists, n-gram databases, and enriched word databases — are sold as one-time perpetual licenses, starting at EUR 250 for academic/research use and EUR 2,500 for commercial use, with custom quotations for large or specialized enterprise databases. The third layer is bespoke professional services covering terminology extraction, document classification, data mining, information retrieval, and custom corpus generation, sold via direct quotation.

The company sells horizontally to three primary buyer types: (i) software developers and AI/NLP practitioners needing reliable, large-scale language data; (ii) dictionary and language-teaching publishers needing lexical databases and frequency information; and (iii) enterprise teams needing language processing solutions for content management. Distribution is exclusively digital — a self-serve website with sample downloads, free trial accounts, custom-quotation enterprise sales, and presence on LinkedIn, Facebook, X, LinkedIn groups, and Academia.edu. The company operates globally with no geographic restrictions disclosed, is privately held with no institutional funding on record, and maintains a small team (11-50 employees).

Short descriptiontext

Lexical Computing generates and sells large-scale language data assets — word frequency lists, n-gram databases, and enriched lexicons built from proprietary corpora totaling 1 trillion words across 100+ languages — alongside corpus query and dictionary editing tools, serving software developers, dictionary publishers, and AI/NLP customers.

Operating statusenum
Operating
Ownership categoryenum
Headcount rangeband
11–50
akta.pro rankint
HeadquartersPortslade, United Kingdom
HQ citystring
Portslade
HQ countrystring
United Kingdom
HQ regionstring
Europe
Markets served

Serves global market

Keyword5 values
corpus linguistics software, word frequency databases, n-gram datasets, lexicography tools, language data services
Industry2 codes
1Search, Retrieval & Semantic Ranking (BM25/vector, hybrid)
CodeHDAAADABPrimaryYes
2Vocabulary, Flashcards, SRS & Memorization Platforms
CodeEDAFAOAFPrimaryNo
NAICS code2 codes
  • Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services5182
  • Web Search Portals, Libraries, Archives, and Other Information Services519
SIC code2 codes
  • Services-Computer Processing & Data Preparation7374
  • Services-Prepackaged Software7372
Product category
Language Data and Corpus Linguistics Software
GTM motion1 record

Each record includes

Type, Description, Source

Revenue model3 records
1Word Database Sales
TypeOne Time License
Description

One-time purchase of word frequency lists, n-gram databases, and lexicons in multiple languages with pricing based on specifications, corpus size, and intended use

lexicalcomputing.com
2Custom Data Services
TypeProfessional Services
Description

Custom language data extraction, terminology extraction, document classification, and information retrieval solutions tailored to customer requirements

lexicalcomputing.com
3Software Tools
TypeSubscription Recurring
Description

Online tools including Sketch Engine corpus query system, Lexonomy dictionary editor, and OneClick Terms terminology extraction

lexicalcomputing.com
Marketing channels5 records

Each record includes

Title, Type, Stage, Description, Source

Distribution channels4 records

Each record includes

Title, Type, Scope, Target buyer, Description, Source

Cost components5 values
Personnel, Technology or R&D, Infrastructure, Marketing or Sales, Operations
Pricing details2 tiers
1Research/Academic Use - starting price
ModelOne time/ perpetual licenseBilling cadencePay-as-you-go
Notes

EUR 250 starting price for academic/research use

lexicalcomputing.com
2Commercial Use - standard price
ModelOne time/ perpetual licenseBilling cadencePay-as-you-go
Notes

EUR 2,500 for commercial use. Discounts offered for non-academic use when purchasing 3 or more lists.

lexicalcomputing.com
GTM typeB2B
B2B
Offering typeSoftware
Software
Brand1 of 4 records shown
1Sketch Engine
Description

Corpus query and management system for language data analysis

lexicalcomputing.com
+3 more records
Core offering1 text field

Lexical Computing supplies large-scale word databases, lexicons, n-gram databases, and word frequency lists generated from text corpora containing approximately 1 trillion words across 100+ languages. It also provides software tools for corpus querying and management (Sketch Engine), online dictionary editing (Lexonomy), terminology extraction (OneClick Terms), and lexicography training (Lexicom), and offers custom language data and information retrieval services. Its customers are primarily software developers, dictionary and language teaching material publishers, and enterprises needing reliable language data.

Differentiator
Functional benefit
Problem solved
Quantifiable outcome1 value
  • Corpus sizes up to 90 billion words for single languages enabling rare word and n-gram frequency data
Product overview1 text field

Lexical Computing provides a portfolio of language data products and tools built around a unified linguistic analysis platform. The core offerings include Sketch Engine (a corpus query and management system), Lexonomy (an online dictionary editor), OneClick Terms (terminology extraction), and Lexicom (educational courses). Supporting these are downloadable data products including word frequency lists, n-gram databases, and enriched word databases/lexicons generated from text corpora containing over 1 trillion words in 100+ languages. The company also provides solutions in full-text search, terminology extraction, document classification and categorization, data mining, and information retrieval.

Product and service7 records
1Sketch Engine
CategoryCorpus query and management software
Description

Corpus query and management system that allows users to search and analyze text corpora for word frequency data, n-grams, collocations, and other linguistic information. Targeted at software developers, researchers, and language professionals.

2Lexonomy
CategoryDictionary editing software
Description

Online dictionary editor for creating and managing digital dictionaries and lexicons. Targeted at lexicographers, dictionary publishers, and language teaching content creators.

3OneClick Terms
CategoryTerminology extraction software
Description

Terminology extraction tool that automatically identifies and extracts technical terms from text documents. Targeted at enterprise data teams, translators, and terminology managers.

4Lexicom
CategoryEducational course in lexicography and lexical computing
Description

A Course in Lexicography and Lexical Computing providing educational training in dictionary creation and language data processing. Targeted at students and professionals in lexicography and computational linguistics.

5Word Frequency Lists
CategoryLanguage data product
Description

High-quality frequency word lists generated from text corpora, available for multiple languages including English, Spanish, French, Arabic, Russian, Portuguese, and Hindi. Targeted at software developers, NLP teams, lexicographers, and publishers.

6N-gram Databases
CategoryLanguage data product
Description

Bigram, trigram, and larger n-gram databases generated from web corpora, available for multiple languages with millions of unique n-grams and options for enrichment with POS tags, lemmas, and probabilities. Targeted at software developers, NLP teams, and computational linguists.

7Word Databases and Lexicons
CategoryLanguage data product
Description

Large word databases, lexical data, word lists, and lexicons enriched with linguistic data such as synonyms, collocations, example sentences, and morphological information. Targeted at dictionary publishers, language teaching material publishers, and software developers.

Scale indicator11 records

Each record includes

Type, Value, Description, Source

Recent move4 records

Each record includes

Date, Type, Title, Description, Source

Expansion highlight4 records

Each record includes

Type, Description

Peers10 records
TypeOthers
Description

LDC is the leading North American distributor of linguistic corpora, lexicons, and annotations to academic and government buyers; it is a distribution-side peer serving the same corpus linguistics and NLP research customers.

TypeBroad incumbent
Description

Hugging Face operates a broad hub for NLP models and datasets including pre-built corpora and language data; it overlaps with Lexical Computing in NLP data distribution and dictionary/lexicon-adjacent assets, but at vastly larger scale and broader scope.

3CQPweb
TypeDirect peer
Description

CQPweb, from Lancaster University, is a self-hosted corpus query system using the IMS Open Corpus Workbench; it is a direct technical peer for Sketch Engine in academic corpus linguistics deployments.

4AntConc
TypeDirect peer
Description

Laurence Anthony's AntConc is a free, widely used desktop corpus toolkit providing concordances, word lists, and n-gram analysis; it competes head-to-head with Sketch Engine for academic and corpus linguistics users and constrains pricing in that segment.

5LancsBox
TypeDirect peer
Description

LancsBox is Lancaster's modern successor desktop corpus toolkit with corpus building, annotation, and analysis features; it competes with Sketch Engine across the same corpus linguistics workflow.

6WordSmith Tools
TypeDirect peer
Description

Mike Scott's WordSmith Tools is a long-standing paid corpus analysis toolkit offering word frequency lists, concordances, and keyword analytics — a near-direct functional peer to Sketch Engine for the same lexicographer and corpus linguist buyer.

TypeOthers
Description

Common Crawl provides open, petabyte-scale web crawl data that many LLM teams use directly as a training corpus; it is the closest open-data alternative to Lexical Computing's cleaned text corpora and shapes market expectations on data pricing.

TypeBroad incumbent
Description

MemoQ and SDL Trados are leading translation environment tools that include terminology management and translation memory features, overlapping with OneClick Terms and Lexonomy for dictionary and terminology workflows used by translators and language service providers.

9ELRA (European Language Resources Association)
TypeOthers
Description

ELRA / ELDA is the principal European distributor and consortium for language resources including corpora and lexica; it operates in the same buyer ecosystem as Lexical Computing and overlaps on European-language data licensing.

10Voyant Tools
TypeDirect peer
Description

Voyant Tools is a web-based, open-source text analysis environment for corpus reading and exploration; developed in the same digital humanities and corpus linguistics tradition, it overlaps directly with Sketch Engine's web-delivered analytical use cases.

Market position
Strengths5 records

Each record includes

Headline, Details, Source

Weaknesses5 records

Each record includes

Headline, Details, Source

Competitive moat3 records

Each record includes

Type, Details

Key risks6 records

Each record includes

Headline, Details, Source

Key highlights7 records

Each record includes

Headline, Details, Source

Customer concentration

Classification, Details

Named customers3 records

Each record includes

Name, Industry, Type, Use case, Source, UUID

Segment3 records

Each record includes

Title, Type, Primary, Description, Pain point addressed, Use case, Source

Ideal customer profile3 records

Each record includes

Profile, Firmographic size, Sales motion, Sales cycle length, Buying structure, Purchase trigger, Buyer persona, Geography, Industry vertical, Primary use case, Description, Pain points, Evidence proof points, Target buyer

Technology focused
Yes
API detail
Has APIbool
No

Docs URL, Description

AI capability2 records

Each record includes

Type, Description, Source

AI maturity
App detail

Has app

Feature5 records

Each record includes

Title, Differentiator, Description, Source

Core technology
Revenue estimate
Valuation estimate
Number of profiles
No data
No data
Funding overview

Funding stage, Last funding date, Total funding USD

Funding rounds

Each record includes

Round, Amount USD, Date, Pre money valuation, Total investors, Investors, News

Investors

Each record includes

Name, Type, Date of entry, Rounds participated, Website

Funding detail is available on the Subscription and Enterprise plan.Contact sales →

M&A

Each record includes

Name, Acquisition type, Announced date, Completed date, Status, Website, News

Investment

Each record includes

Name, Round, Announced date, Lead investor, Website, News

M&A and investment is available on the Subscription and Enterprise plan.Contact sales →

Lexical Computing

Language Data and Corpus Linguistics Softwarelexicalcomputing.com

Lexical Computing generates and sells large-scale language data assets — word frequency lists, n-gram databases, and enriched lexicons built from proprietary corpora totaling 1 trillion words across 100+ languages — alongside corpus query and dictionary editing tools, serving software developers, dictionary publishers, and AI/NLP customers.

What Lexical Computing does

Lexical Computing, founded in 2003 and headquartered in Portslade, United Kingdom (with an operating entity in the Czech Republic), builds and sells large-scale linguistic data assets and the tooling to use them. Its core technology is a proprietary pipeline that collects and cleans authentic web text to construct text corpora totaling 1 trillion words across 100+ languages, with the largest English corpus at 90 billion words. From these corpora the company generates enriched products — word frequency lists, n-gram (bigram, trigram and larger) databases, and word databases/lexicons containing millions of unique items per language with POS tags, lemmas, probabilities, synonyms, collocations, example sentences, and morphological information. The company positions these datasets as inputs for language modeling (LLMs), NLP applications, and corpus-based research.

The commercial portfolio spans three categories. Self-serve software tools — Sketch Engine (corpus query and management), Lexonomy (online dictionary editor), OneClick Terms (terminology extraction), and Lexicom (a course in lexicography) — operate on a subscription/recurring basis with free trials. Downloadable data products — word frequency lists, n-gram databases, and enriched word databases — are sold as one-time perpetual licenses, starting at EUR 250 for academic/research use and EUR 2,500 for commercial use, with custom quotations for large or specialized enterprise databases. The third layer is bespoke professional services covering terminology extraction, document classification, data mining, information retrieval, and custom corpus generation, sold via direct quotation.

The company sells horizontally to three primary buyer types: (i) software developers and AI/NLP practitioners needing reliable, large-scale language data; (ii) dictionary and language-teaching publishers needing lexical databases and frequency information; and (iii) enterprise teams needing language processing solutions for content management. Distribution is exclusively digital — a self-serve website with sample downloads, free trial accounts, custom-quotation enterprise sales, and presence on LinkedIn, Facebook, X, LinkedIn groups, and Academia.edu. The company operates globally with no geographic restrictions disclosed, is privately held with no institutional funding on record, and maintains a small team (11-50 employees).

Lexical Computing firmographics

Firmographics
Name
Lexical Computing
Legal name
Lexical Computing
Website
https://lexicalcomputing.com
Company type
Private
Founded year
2003
Operating status
Operating
Headcount range
11–50 employees
Short description
Lexical Computing generates and sells large-scale language data assets — word frequency lists, n-gram databases, and enriched lexicons built from proprietary corpora totaling 1 trillion words across 100+ languages — alongside corpus query and dictionary editing tools, serving software developers, dictionary publishers, and AI/NLP customers.
Ownership category
akta.pro rank

Lexical Computing industry classification

Industry
Product category
Language Data and Corpus Linguistics Software
NAICS
Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (5182), Web Search Portals, Libraries, Archives, and Other Information Services (519)
SIC
Services-Computer Processing & Data Preparation (7374), Services-Prepackaged Software (7372)
akta.pro primary industry
Search, Retrieval & Semantic Ranking (BM25/vector, hybrid) (HDAAADAB)
akta.pro secondary industry
Vocabulary, Flashcards, SRS & Memorization Platforms (EDAFAOAF)

Keywords

  • Corpus linguistics software
  • Word frequency databases
  • N-gram datasets
  • Lexicography tools
  • Language data services

Where Lexical Computing is headquartered

Location

Headquarters

HQ city
Portslade
HQ country
United Kingdom
HQ region
Europe

Markets served

Lexical Computing business model

Business model
GTM type
B2B
Offering type
Software
Cost components
Personnel, Technology or R&D, Infrastructure, Marketing or Sales, Operations

Revenue model

  1. Word Database Sales: One-time purchase of word frequency lists, n-gram databases, and lexicons in multiple languages with pricing based on specifications, corpus size, and intended use
  2. Custom Data Services: Custom language data extraction, terminology extraction, document classification, and information retrieval solutions tailored to customer requirements
  3. Software Tools: Online tools including Sketch Engine corpus query system, Lexonomy dictionary editor, and OneClick Terms terminology extraction

Pricing tiers

ModelBillingPrice
One time/ perpetual licensePay-as-you-goResearch/Academic Use - starting price
One time/ perpetual licensePay-as-you-goCommercial Use - standard price

Go-to-market motion1 record

Distribution channels4 records

Marketing channels5 records

Lexical Computing product offering

Product offering

Core offering

Lexical Computing supplies large-scale word databases, lexicons, n-gram databases, and word frequency lists generated from text corpora containing approximately 1 trillion words across 100+ languages. It also provides software tools for corpus querying and management (Sketch Engine), online dictionary editing (Lexonomy), terminology extraction (OneClick Terms), and lexicography training (Lexicom), and offers custom language data and information retrieval services. Its customers are primarily software developers, dictionary and language teaching material publishers, and enterprises needing reliable language data.

Product overview

Lexical Computing provides a portfolio of language data products and tools built around a unified linguistic analysis platform. The core offerings include Sketch Engine (a corpus query and management system), Lexonomy (an online dictionary editor), OneClick Terms (terminology extraction), and Lexicom (educational courses). Supporting these are downloadable data products including word frequency lists, n-gram databases, and enriched word databases/lexicons generated from text corpora containing over 1 trillion words in 100+ languages. The company also provides solutions in full-text search, terminology extraction, document classification and categorization, data mining, and information retrieval.

Differentiator

Problem solved

Functional benefit

Brands

  • Sketch Engine: Corpus query and management system for language data analysis
  • Lexonomy
  • OneClick Terms
  • Lexicom

Products and services

  • Sketch Engine Corpus query and management system that allows users to search and analyze text corpora for word frequency data, n-grams, collocations, and other linguistic information. Targeted at software developers, researchers, and language professionals.
  • Lexonomy Online dictionary editor for creating and managing digital dictionaries and lexicons. Targeted at lexicographers, dictionary publishers, and language teaching content creators.
  • OneClick Terms Terminology extraction tool that automatically identifies and extracts technical terms from text documents. Targeted at enterprise data teams, translators, and terminology managers.
  • Lexicom A Course in Lexicography and Lexical Computing providing educational training in dictionary creation and language data processing. Targeted at students and professionals in lexicography and computational linguistics.
  • Word Frequency Lists High-quality frequency word lists generated from text corpora, available for multiple languages including English, Spanish, French, Arabic, Russian, Portuguese, and Hindi. Targeted at software developers, NLP teams, lexicographers, and publishers.
  • N-gram Databases Bigram, trigram, and larger n-gram databases generated from web corpora, available for multiple languages with millions of unique n-grams and options for enrichment with POS tags, lemmas, and probabilities. Targeted at software developers, NLP teams, and computational linguists.
  • Word Databases and Lexicons Large word databases, lexical data, word lists, and lexicons enriched with linguistic data such as synonyms, collocations, example sentences, and morphological information. Targeted at dictionary publishers, language teaching material publishers, and software developers.

Quantifiable outcome

  • Corpus sizes up to 90 billion words for single languages enabling rare word and n-gram frequency data

Companies that use Lexical Computing

Customer profile

Named customers3 records

Segments3 records

Ideal customer profiles3 records

Lexical Computing technology and API

Technology

Technology focussed Yes

API detail

Has API
No
API docs
API detail

Core technology

AI maturity

App detail

AI capability2 records

Feature5 records

Lexical Computing partnerships and signals

Strategic signal

Scale indicators11 records

Recent moves4 records

Expansion highlights4 records

Lexical Computing competitors and assessment

Company assessment

Others

  • Linguistic Data Consortium (LDC): LDC is the leading North American distributor of linguistic corpora, lexicons, and annotations to academic and government buyers; it is a distribution-side peer serving the same corpus linguistics and NLP research customers.
  • Common Crawl Foundation: Common Crawl provides open, petabyte-scale web crawl data that many LLM teams use directly as a training corpus; it is the closest open-data alternative to Lexical Computing's cleaned text corpora and shapes market expectations on data pricing.
  • ELRA (European Language Resources Association): ELRA / ELDA is the principal European distributor and consortium for language resources including corpora and lexica; it operates in the same buyer ecosystem as Lexical Computing and overlaps on European-language data licensing.

Broad incumbents

  • Hugging Face: Hugging Face operates a broad hub for NLP models and datasets including pre-built corpora and language data; it overlaps with Lexical Computing in NLP data distribution and dictionary/lexicon-adjacent assets, but at vastly larger scale and broader scope.
  • MemoQ / SDL Trados: MemoQ and SDL Trados are leading translation environment tools that include terminology management and translation memory features, overlapping with OneClick Terms and Lexonomy for dictionary and terminology workflows used by translators and language service providers.

Direct peers

  • CQPweb: CQPweb, from Lancaster University, is a self-hosted corpus query system using the IMS Open Corpus Workbench; it is a direct technical peer for Sketch Engine in academic corpus linguistics deployments.
  • AntConc: Laurence Anthony's AntConc is a free, widely used desktop corpus toolkit providing concordances, word lists, and n-gram analysis; it competes head-to-head with Sketch Engine for academic and corpus linguistics users and constrains pricing in that segment.
  • LancsBox: LancsBox is Lancaster's modern successor desktop corpus toolkit with corpus building, annotation, and analysis features; it competes with Sketch Engine across the same corpus linguistics workflow.
  • WordSmith Tools: Mike Scott's WordSmith Tools is a long-standing paid corpus analysis toolkit offering word frequency lists, concordances, and keyword analytics — a near-direct functional peer to Sketch Engine for the same lexicographer and corpus linguist buyer.
  • Voyant Tools: Voyant Tools is a web-based, open-source text analysis environment for corpus reading and exploration; developed in the same digital humanities and corpus linguistics tradition, it overlaps directly with Sketch Engine's web-delivered analytical use cases.

Market position

Strengths5 records

Weaknesses5 records

Competitive moat3 records

Key risks6 records

Key highlights7 records

Customer concentration

Lexical Computing social profiles

Digital presence

Lexical Computing financial estimates

Financial estimate

Revenue estimate

Valuation estimate

Lexical Computing leadership team

Management profile

Number of profiles

Lexical Computing funding detail

Funding detail

Funding overview

Funding rounds

Investors

Funding detail is available on the Subscription and Enterprise plan.Contact sales →

Lexical Computing M&A and investment

M&A and investment

M&A

Investments

M&A and investment is available on the Subscription and Enterprise plan.Contact sales →

Frequently asked questions about Lexical Computing

What does Lexical Computing do?

Lexical Computing supplies large-scale word databases, lexicons, n-gram databases, and word frequency lists generated from text corpora containing approximately 1 trillion words across 100+ languages. It also provides software tools for corpus querying and management (Sketch Engine), online dictionary editing (Lexonomy), terminology extraction (OneClick Terms), and lexicography training (Lexicom), and offers custom language data and information retrieval services. Its customers are primarily software developers, dictionary and language teaching material publishers, and enterprises needing reliable language data.

Is Lexical Computing a public or private company?

Lexical Computing is a private company. It is classified as unknown and is currently operating.

When was Lexical Computing founded?

Lexical Computing was founded in 2003. It employs 11 to 50 people.

Where is Lexical Computing based?

Lexical Computing is headquartered in Portslade, United Kingdom, in the Europe region.

How does Lexical Computing make money?

Three revenue lines are on record. Word Database Sales are the primary driver. The others are custom Data Services and software Tools.

Who are Lexical Computing's main competitors?

Others on record are Linguistic Data Consortium (LDC), Common Crawl Foundation and ELRA (European Language Resources Association). Broad incumbents are Hugging Face and MemoQ / SDL Trados. Direct peers are CQPweb, AntConc, LancsBox, WordSmith Tools and Voyant Tools.

Does Lexical Computing have an API?

No public API is recorded for Lexical Computing.

What industry is Lexical Computing in?

Lexical Computing's product category is Language Data and Corpus Linguistics Software. Its primary akta.pro industry code is HDAAADAB, Search, Retrieval & Semantic Ranking (BM25/vector, hybrid), with a secondary code of EDAFAOAF, Vocabulary, Flashcards, SRS & Memorization Platforms. Its NAICS code is 5182 and its SIC code is 7374.

Unlock the full company data

50 free credits on sign-up, no credit card required.

Contact sales