Doubleword
Doubleword is a London-based AI infrastructure company providing an OpenAI-compatible API for high-throughput asynchronous and batch LLM inference, delivering open-weight model serving at up to 80-90% lower cost than real-time frontier APIs for AI engineering teams running background agents, data pipelines, evaluations, and synthetic data generation.
- Company typePrivate
- Founded2021
- HeadquartersLondon, United Kingdom
- Headcount11–50
- GTM typeB2B
- OfferingSoftware
What Doubleword does
Doubleword is a London-based AI infrastructure company (legally TYTN LTD) that provides an OpenAI-compatible API for high-throughput asynchronous and batch LLM inference, targeting AI engineering teams whose unit economics are broken by real-time API pricing on background workloads. The platform is built on a proprietary five-layer inference stack consisting of a Rust-based model gateway (Control Layer), intelligent GPU scheduling and orchestration, an optimized runtime engine built on TensorRT-LLM, SGLang, and vLLM with custom techniques including speculative KV coding, Cloudburst cold-start acceleration, and ZeroDP JIT weight offloading, a curated portfolio of open-weight models (DeepSeek, Qwen, Kimi, GLM, Gemma, Nemotron families across text, vision, OCR, and embeddings), and flexible multi-cloud GPU infrastructure that uses spot instances to minimize cost. Customers access three pricing tiers (Realtime, Async at 25-50% off, Batch/Overnight at 50-80% off), with documented cost outcomes of 4-9x below frontier closed APIs on comparable-intelligence workloads.
The company operates a product-led, self-serve go-to-market: developers sign up at app.doubleword.ai, receive an instant API key with no credit card required, and migrate by changing the base URL on existing OpenAI client code. Supporting tooling includes open-source libraries (Autobatcher for transparent request batching, QLM query language), production-ready workbooks on GitHub covering async agents, synthetic data, evals, and ETL pipelines, and named integrations with developer platforms (Arize AI, LangSmith, OpenClaw). Reference customers include OpenMed (medical VQA at 94% below Anthropic), Dataiku's 575 Lab (synthetic PII generation at $50 total cost), and a community of individual developers publishing use cases on LinkedIn, Medium, and Hugging Face.
Doubleword raised a $12 million seed round led by Dawn Capital in May 2025 (with prior funding from Octopus Ventures and Intel Ignite), and in April 2026 was selected as one of six startups to receive up to one million GPU hours on the UK's AI Research Resource supercomputer network under the £500M Sovereign AI initiative. The company is privately held, headquartered in London, and structured around an inference-systems engineering team of 11-50 employees.
Doubleword firmographics
Firmographics- Name
- Doubleword
- Legal name
- TYTN LTD
- Website
- https://doubleword.ai
- Company type
- Private
- Founded year
- 2021
- Operating status
- Operating
- Headcount range
- 11–50 employees
- Short description
- Doubleword is a London-based AI infrastructure company providing an OpenAI-compatible API for high-throughput asynchronous and batch LLM inference, delivering open-weight model serving at up to 80-90% lower cost than real-time frontier APIs for AI engineering teams running background agents, data pipelines, evaluations, and synthetic data generation.
- Ownership category
- akta.pro rank
Doubleword industry classification
Industry- Product category
- Cloud AI Inference
- NAICS
- Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (518), Custom Computer Programming Services (541511), Computer Systems Design and Related Services (54151)
- SIC
- Services-Computer Programming, Data Processing, Etc. (7370), Services-Computer Integrated Systems Design (7373), Services-Computer Programming Services (7371)
- akta.pro primary industry
- Model Hosting, Serving & Inference Platforms (HDAAACAB)
- akta.pro secondary industries
- Model Deployment, Serving & Inference Platforms (HDAAABAF), Model Serving, Inference & Deployment Platforms (APIs, Edge/On-Prem) (HDAEANAC), Model Compression & Optimization (Quantization, Distillation, Pruning) (HDAAACAL), Generative AI & LLM Solutions Services (RAG, Agents, Copilots) (BPAEAHAG)
Keywords
Where Doubleword is headquartered
LocationHeadquarters
- HQ city
- London
- HQ country
- United Kingdom
- HQ region
- Europe
Offices1 record
Markets served
Doubleword business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Infrastructure, Technology or R&D, Personnel, Marketing or Sales, Operations
Revenue model
- LLM Inference API (Usage-Based): Pay-per-token pricing for API access to hosted open-weight LLM models. Three tiers: Realtime (full price, latency-optimized), Async (25-50% discount, high-throughput), and Batch/Overnight (50-80% discount, 24H SLA). No minimum spend, no credit card required. Revenue scales with token volume consumed by customers.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Usage-based | Pay-as-you-go | Realtime API — full price, latency-optimized |
| Usage-based | Pay-as-you-go | Async Inference — 25-50% off realtime |
| Usage-based | Pay-as-you-go | Batch / Overnight — 50-80% off realtime, 24H SLA |
Go-to-market motion2 records
Distribution channels3 records
Marketing channels8 records
Doubleword product offering
Product offeringCore offering
Doubleword operates a high-throughput AI inference API platform delivering hosted access to open-weight large language models (LLMs) at up to 80–90% lower cost than real-time frontier APIs. The platform offers three pricing tiers (Realtime, Async, Batch/Overnight) backed by a proprietary five-layer inference stack — a Rust-based model gateway, intelligent orchestration, an optimized runtime engine (TensorRT-LLM/SGLang/vLLM), curated quantized open-weight models, and flexible multi-cloud GPU infrastructure.
Product overview
Doubleword is a London-based AI inference provider offering a unified platform for high-throughput async and batch inference at up to 90% lower cost than real-time APIs. The core product is the Doubleword API (OpenAI-compatible at https://api.doubleword.ai/v1), which delivers three pricing tiers: Realtime, Async, and Batch/Overnight. This is powered by the Doubleword Inference Stack—a proprietary five-layer infrastructure (gateway, orchestration, runtime engine, models, hardware). The platform provides access to an extensive portfolio of open-weight models including text models (DeepSeek-V4-Pro, Qwen3.5-397B, GLM-5.2, Kimi-K2.6), vision-language models (Qwen3-VL), embedding models (Qwen3-Embedding-8B), and OCR models (DeepSeek-OCR-2, olmOCR-2-7B, LightOnOCR-2-1B). Supporting resources include Workbooks (production-ready templates), the Autobatcher library (drop-in batching), the open-source Control Layer gateway, a Savings Calculator, and an AI Glossary. Integrations with Arize AI, LangSmith, and OpenClaw extend evaluation and agent workflow capabilities.
Differentiator
Problem solved
Functional benefit
Products and services
- Doubleword Inference API Public OpenAI-compatible REST API for high-throughput async and batch inference of open-weight LLMs. Supports Chat Completions, Responses, Embeddings, and Batch endpoints with tool calling, structured JSON generation, and webhook delivery. Three pricing tiers (Realtime, Async 25-50% off, Batch 50-80% off) targeting developers and engineering teams running background agents, evaluations, ETL pipelines, and synthetic data generation.
- Open-Weight Models Portfolio Catalog of open-weight LLMs hosted and served through the Doubleword API, including text generation models (DeepSeek-V4-Pro/Flash, Qwen3.5 family, GLM-5.2, Kimi-K2.6, Gemma-4-31B, Nemotron-3-Ultra-550B, GPT-OSS-20B), vision-language models (Qwen3-VL-235B/30B), embedding models (Qwen3-Embedding-8B), and OCR/document models (DeepSeek-OCR-2, olmOCR-2-7B, LightOnOCR-2-1B). Models are delivered with INT8/FP8/INT4 quantization for cost efficiency.
- Autobatcher Open-source Python library acting as a drop-in AsyncOpenAI replacement that transparently batches requests over a configurable window to capture batch-tier pricing without code rewrites or .jsonl file management. Enables existing async OpenAI code to switch to Doubleword by changing the base URL.
- Control Layer Open-source Rust-based model gateway handling multi-model routing, access controls, logging, and monitoring at scale. Claims 450× less overhead than LiteLLM. Designed for teams self-hosting LLM routing infrastructure.
- QLM (Query Language for Models) Open-source query language tool providing a structured interface for interacting with models hosted on Doubleword's API. Available on GitHub.
Quantifiable outcome
- Up to 90% cost reduction vs Anthropic Claude Opus 4.6 real-time API for comparable intelligence (Qwen3.5-4B Batch: $100 vs Anthropic ~$30K per 1B in+out)
- +8 more outcomes
Companies that use Doubleword
Customer profileNamed customers6 records
Segments4 records
Ideal customer profiles4 records
Doubleword technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration3 records
AI capability9 records
Feature11 records
Doubleword partnerships and signals
Strategic signalPartnerships
Six partnerships are on record, tiered core and secondary.
- Arize AIcoreIntegration between Doubleword's batch inference API and Arize AX observability platform. Doubleword configured as a Custom Model Endpoint in Arize AX, enabling continuous LLM-as-judge evaluations with OpenTelemetry trace visualization. Frontier judge models (DeepSeek V4 Pro) run at batch cost; Arize handles trace topology and scoring dashboards.
- LangSmithsecondaryLangSmith integration allowing users to run LLM-as-judge evaluations via Doubleword's batch API. Demonstrated $8.82 for 10K evals — 97% cheaper than GPT-5.5 real-time with zero 429 errors.
- OpenClawsecondaryOpenClaw computer-use agent platform has an official Doubleword skill enabling background inference for always-on agents. Agents route background tasks (research, email triage, doc processing, cron jobs) to Doubleword async tier at up to 10× lower cost than real-time endpoints. Skill published at github.com/doublewordai/batch-skill.
- OpenMedsecondaryOpenMed used Doubleword's batch inference API for SynthVision project — 119,137 medical images annotated using Qwen 397B and Kimi K2.5 frontier VLMs at $452.58 total. Resulting 110K medical VQA records open-sourced on HuggingFace. Doubleword serves as the inference infrastructure partner for the research collaboration.
- DataikusecondaryDataiku's 575 Lab used Doubleword to generate synthetic PII training data for Kiji Privacy Proxy — a privacy protection model for enterprise AI workflows. $50 cost for 26 PII entity types vs ~$1,000 with closed-source providers. Kiji Privacy Proxy and training dataset published under Apache 2.0 on HuggingFace.
- Hugging FacesecondaryHugging Face hosts the SynthVision dataset (OpenMed × HuggingFace collaboration) — 110K synthetic medical VQA records produced using Doubleword's batch inference. The dataset and resulting model adapters are open-sourced on HuggingFace's platform, with Doubleword credited as the inference infrastructure.
Scale indicators3 records
Recent moves6 records
Expansion highlights6 records
Doubleword competitors and assessment
Company assessmentBroad incumbents
- AWS Bedrock: AWS Bedrock is a managed service offering foundation models including open-weight options, bundled with broader AWS infrastructure and enterprise SLAs. It competes with Doubleword on open-weight inference at scale but as part of a much larger cloud portfolio.
- Google Vertex AI: Vertex AI is Google's enterprise AI platform offering model hosting, fine-tuning, and inference including open-weight models. It competes broadly with Doubleword in serving enterprise LLM workloads, particularly for organizations already committed to Google Cloud.
- Azure AI Foundry: Azure AI Foundry (formerly Azure AI Studio) hosts open-weight and proprietary models with enterprise-grade governance and procurement integration. It is a broad incumbent competing with Doubleword for enterprise inference workloads.
Direct peers
- DeepInfra: DeepInfra is a low-cost serverless inference provider for open-weight models, explicitly positioning on price-per-token economics. It is a direct functional competitor to Doubleword's async and batch tiers.
- Together AI: Together AI is a direct competitor offering a cloud platform for open-weight LLM inference and fine-tuning at low per-token pricing. Both target developers seeking cheap, high-throughput access to open-weight models and emphasize cost-per-token as the primary differentiator.
- Replicate: Replicate hosts open-weight AI models via a pay-per-prediction API with strong developer ergonomics. It is comparable to Doubleword in target customer (developers running open models) and pricing model (usage-based), though with broader model coverage beyond LLMs.
- Anyscale: Anyscale (built on Ray) offers managed compute and inference for AI workloads, including open-weight LLM serving. It overlaps with Doubleword's focus on efficient inference at scale and serves similar ML engineering audiences.
- Fireworks AI: Fireworks AI provides a serverless inference platform for open-weight LLMs and custom models, with a strong focus on throughput and low latency. It competes head-to-head with Doubleword on cost-per-token and supports async/batch workloads.
Emerging players
- Modal Labs: Modal provides serverless compute for AI and data workloads, including LLM inference. It overlaps with Doubleword in serving developers running GPU-accelerated jobs but with a more general-purpose compute framing.
- Lepton AI: Lepton AI offers a cloud platform for running and deploying open-weight AI models, with emphasis on developer experience and cost efficiency. It targets similar ML engineering buyers as Doubleword.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat5 records
Key risks6 records
Key highlights7 records
Customer concentration
Doubleword social profiles
Digital presenceDoubleword financial estimates
Financial estimateRevenue estimate
Valuation estimate
Doubleword leadership team
Management profileNumber of profiles
Doubleword funding detail
Funding detailFunding overview
Funding rounds3 records
Investors4 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Doubleword M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Doubleword
What does Doubleword do?
Doubleword operates a high-throughput AI inference API platform delivering hosted access to open-weight large language models (LLMs) at up to 80–90% lower cost than real-time frontier APIs. The platform offers three pricing tiers (Realtime, Async, Batch/Overnight) backed by a proprietary five-layer inference stack — a Rust-based model gateway, intelligent orchestration, an optimized runtime engine (TensorRT-LLM/SGLang/vLLM), curated quantized open-weight models, and flexible multi-cloud GPU infrastructure.
Is Doubleword a public or private company?
Doubleword is a private company. It is classified as venture growth investor backed and is currently operating.
When was Doubleword founded?
Doubleword was founded in 2021. It employs 11 to 50 people.
Where is Doubleword based?
Doubleword is headquartered in London, United Kingdom, in the Europe region.
How does Doubleword make money?
One revenue line is on record: LLM Inference API (Usage-Based).
Who are Doubleword's main competitors?
Broad incumbents on record are AWS Bedrock, Google Vertex AI and Azure AI Foundry. Direct peers are DeepInfra, Together AI, Replicate, Anyscale and Fireworks AI. Emerging players are Modal Labs and Lepton AI.
Does Doubleword have an API?
Yes. Doubleword offers a public OpenAI-compatible REST API at https://api.doubleword.ai/v1 for high-throughput async and batch inference. The API supports Chat Completions, Responses, Embeddings, and Batch endpoints. It provides full tool calling and structured generation support. Developers can migrate from OpenAI by simply changing the base URL. The API offers three pricing tiers: Realtime (latency optimized), Async (high throughput, 25-50% off realtime), and Batch/Overnight (24H SLA, 50-80% off realtime). Webhooks are used to notify when async jobs complete. Developer documentation is at docs.doubleword.ai/batches.
What industry is Doubleword in?
Doubleword's product category is Cloud AI Inference. Its primary akta.pro industry code is HDAAACAB, Model Hosting, Serving & Inference Platforms, with a secondary code of HDAAABAF, Model Deployment, Serving & Inference Platforms. Its NAICS code is 518 and its SIC code is 7370.