Kog
Kog is a Paris-based AI infrastructure startup (founded 2023) building a hardware-software co-designed inference engine that delivers 3,000 tokens/s per request on AMD MI300X GPUs for AI coding agents and agentic workflows, via a proprietary monokernel runtime, custom KCCL communication library, and Laneformer architecture.
- Company typePrivate
- Founded2023
- HeadquartersParis, France
- Headcount11–50
- GTM typeB2B
- OfferingSoftware
What Kog does
Kog (legal name: Kog Labs; DBA: Kog AI) is a Paris-based AI infrastructure startup founded in 2023 by Gaël Delalleau, an École Polytechnique engineer with a background in cybersecurity research and high-performance GPU work. The company builds a hardware-software co-designed inference engine — the Kog Inference Engine (KIE) — purpose-built to eliminate software-imposed speed ceilings on standard datacenter GPUs. KIE integrates a proprietary Laneformer Transformer architecture featuring Delayed Tensor Parallelism (DTP), a single-kernel 'monokernel' runtime, the custom Kog Communication Library (KCCL), and topology-aware GPU memory access; all components are written from scratch in low-level CUDA/HIP and assembly. The system achieves 3,000 output tokens/s per request on 8× AMD MI300X and 2,100 on 8× NVIDIA H200 at batch size 1 in FP16, without speculative decoding, and is exposed as a drop-in replacement for vLLM.
Kog monetizes via API access sold to teams building AI coding agents and agentic workflows, supplemented by a high-touch Design Partner Program for enterprise engagements; no public pricing tiers are disclosed. The company also publishes the Laneformer 2B (2.3B-parameter) coding model on Hugging Face as both a usable checkpoint and a developer-acquisition channel. Distribution is API-first with a live public playground (playground.kog.ai), a technical blog (blog.kog.ai), and featured presence at AMD AI DevDay 2026. The company is backed by Varsity VC and BPI France's Deep Tech Program ($5M seed) and was awarded the French Tech 2030 label by the French government in October 2025; current operations are anchored by an 11-person team.
Kog firmographics
Firmographics- Name
- Kog
- Legal name
- Kog Labs
- Website
- https://kog.ai
- Company type
- Private
- Founded year
- 2023
- Operating status
- Operating
- Headcount range
- 11–50 employees
- Short description
- Kog is a Paris-based AI infrastructure startup (founded 2023) building a hardware-software co-designed inference engine that delivers 3,000 tokens/s per request on AMD MI300X GPUs for AI coding agents and agentic workflows, via a proprietary monokernel runtime, custom KCCL communication library, and Laneformer architecture.
- Ownership category
- akta.pro rank
Kog industry classification
Industry- Product category
- AI Inference Infrastructure
- NAICS
- Software Publishers (5132), Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (51821), Custom Computer Programming Services (541511)
- SIC
- Services-Prepackaged Software (7372)
- akta.pro primary industry
- Model Deployment, Serving & Inference Platforms (HDAAABAF)
- akta.pro secondary industries
- AI Compiler, Runtime & Kernel Optimization Software (CUDA/ROCm/XLA, graph compilers) (HDAAAAAI), Model Serving, Inference & Deployment Platforms (APIs, Edge/On-Prem) (HDAEANAC), Model Hosting, Serving & Inference Platforms (HDAAACAB)
Keywords
Where Kog is headquartered
LocationHeadquarters
- HQ city
- Paris
- HQ country
- France
- HQ region
- Europe
Offices1 record
Markets served
Kog business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Personnel
Revenue model
- API Access / Inference-as-a-Service: Kog likely generates revenue by providing API access to its high-speed inference engine. The website prominently features 'Request API Access' CTAs and a public tech preview playground, suggesting a consumption-based or subscription API pricing model targeting developers and enterprises building AI agents. No specific pricing tiers are publicly disclosed.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Usage-based | Pay-as-you-go | API access for developers and enterprises |
Go-to-market motion4 records
Distribution channels3 records
Marketing channels6 records
Kog product offering
Product offeringCore offering
Kog builds a hardware-software co-designed inference engine (KIE) for ultra-low-latency LLM inference on standard datacenter GPUs (AMD MI300X and NVIDIA H200), achieving 3,000 output tokens per second per request. The platform combines a custom Laneformer model architecture with Delayed Tensor Parallelism (DTP), a monokernel GPU runtime, and a custom KCCL communication library, exposing inference through a vLLM-compatible API and a live playground. The offering targets enterprises building AI coding agents and agentic workflows where token generation speed governs agent iteration throughput.
Product overview
Kog is a Paris-based AI infrastructure startup offering a unified inference platform centered on the Kog Inference Engine (KIE), a hardware-software co-designed system for ultra-low-latency LLM inference. The core product integrates the Laneformer model architecture with Delayed Tensor Parallelism (DTP), the custom Kog Communication Library (KCCL), and a monokernel runtime. The platform includes Kog Labs for research/blog content and a live Playground demo. Users can access the inference engine via API request or through the drop-in vLLM-compatible interface, targeting AI coding agents and agentic workflows that require real-time token generation speeds of 3,000 tokens/s per request.
Differentiator
Problem solved
Functional benefit
Brands
- Kog Inference Engine (KIE): Hardware-software co-design inference engine enabling 3,000 tokens/s per request on AMD MI300X GPUs.
- Laneformer
- Kog CommunicationLibrary (KCCL)
- Kog Playground
Products and services
- Kog Inference Engine (KIE) A hardware-software co-designed inference engine built to run GPUs at their absolute ceiling, delivering 3,000 tokens/s per request via a monokernel runtime, custom KCCL communication library, and Laneformer architecture. Functions as a drop-in vLLM replacement with no code refactoring required, targeting enterprise teams building AI coding agents and agentic workflows.
- Laneformer 2B A 2.3B-parameter instruction-tuned coding model built around Kog's Delayed Tensor Parallelism architecture, designed for AI coding agents and software engineering workflows. Achieves 45.1% on HumanEval+ and 51.6% on MBPP+ while running at 3,000 tokens/s on AMD MI300X.
- Kog Laneformer Architecture A novel Transformer architecture variant in which inter-device communication is delayed by one layer so that compute runs continuously without synchronization pauses, enabling linear scaling across high-end GPUs for low-latency inference.
- Kog Communication Library (KCCL) A custom collective communication layer that replaces standard NCCL/RCCL to unlock linear scaling for tensor parallelism across high-end GPUs, achieving sub-3 microsecond AllReduce latency on AMD MI300X.
Quantifiable outcome
- 3,000 output tokens/s/request on AMD MI300X and 2,100 on NVIDIA H200 at batch size 1, FP16, no speculative decoding
- +6 more outcomes
Companies that use Kog
Customer profileSegments2 records
Ideal customer profiles2 records
Kog technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- No
- API docs
- API detail
Core technology
AI maturity
App detail
Integration2 records
AI capability5 records
Feature8 records
Kog partnerships and signals
Strategic signalPartnerships
One partnership is on record.
- AMDcoreAMD featured Kog as an ecosystem partner at AMD AI DevDay 2026 and published Kog's inference benchmark results on AMD's official engineering blog, validating Kog's 3.5x speed improvement on AMD Instinct MI300X GPUs. AMD provides the primary hardware platform for Kog's performance differentiation.
Scale indicators6 records
Recent moves6 records
Expansion highlights6 records
Kog competitors and assessment
Company assessmentDirect peers
- Modal Labs: Modal provides serverless GPU compute and inference infrastructure for AI applications with a developer-first, Python-native model. Comparable in targeting AI agent builders with low-friction GPU access and inference APIs.
- Replicate: Replicate runs a cloud API for running open-source ML models, including LLM inference, with a focus on simplicity and developer ergonomics. Comparable in API-first inference distribution to developers building AI applications.
- Groq: Groq sells LPU-based ultra-low-latency inference-as-a-service, directly competing on tokens-per-second-per-request as the primary value proposition. Targets developer and enterprise agentic workloads with similar speed-first positioning.
- DeepInfra: DeepInfra provides serverless inference APIs for open-source LLMs with custom inference engine optimizations. Directly comparable API-first inference-as-a-service business model targeting developer adoption.
- Together AI: Together AI runs an inference cloud with proprietary optimizations on multi-GPU stacks, including custom kernels and inference engine work. Comparable as an open-model inference API competitor targeting developer and enterprise agent builders.
- Lepton AI: Lepton AI offers a cloud platform for running AI models with custom inference optimizations on multi-GPU hardware. Comparable as a developer-facing inference platform emphasizing performance and flexible deployment.
- Anyscale: Anyscale (Ray) offers distributed compute and serving infrastructure optimized for AI workloads, including LLM inference. Comparable as a developer-focused platform for scaling and serving large models on multi-GPU clusters.
- Fireworks AI: Fireworks AI operates a model inference and fine-tuning platform optimized for low-latency serving of open-source LLMs. Directly comparable in API-first inference distribution and agent workload targeting.
Broad incumbents
- Hugging Face: Hugging Face operates Inference Endpoints and text-generation-inference (TGI), serving open-source models via API. Comparable as an inference platform serving the same open-model ecosystem, but as a broad incumbent rather than a speed-specialized player.
Emerging players
- Cerebras Systems: Cerebras builds purpose-built inference silicon (CS-3/WSE) and a cloud inference service claiming class-leading tokens/second for LLM serving. Comparable in the speed-first inference narrative, but pursues a custom-silicon route rather than Kog's software-only approach on commodity GPUs.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat5 records
Key risks6 records
Key highlights7 records
Customer concentration
Kog social profiles
Digital presenceKog financial estimates
Financial estimateRevenue estimate
Valuation estimate
Kog leadership team
Management profileNumber of profiles
Profiles1 record
Kog funding detail
Funding detailFunding overview
Funding rounds1 record
Investors1 record
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Kog M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Kog
What does Kog do?
Kog builds a hardware-software co-designed inference engine (KIE) for ultra-low-latency LLM inference on standard datacenter GPUs (AMD MI300X and NVIDIA H200), achieving 3,000 output tokens per second per request. The platform combines a custom Laneformer model architecture with Delayed Tensor Parallelism (DTP), a monokernel GPU runtime, and a custom KCCL communication library, exposing inference through a vLLM-compatible API and a live playground. The offering targets enterprises building AI coding agents and agentic workflows where token generation speed governs agent iteration throughput.
Is Kog a public or private company?
Kog is a private company. It is classified as venture growth investor backed and is currently operating.
When was Kog founded?
Kog was founded in 2023. It employs 11 to 50 people.
Where is Kog based?
Kog is headquartered in Paris, France, in the Europe region.
How does Kog make money?
One revenue line is on record: API Access / Inference-as-a-Service.
Who are Kog's main competitors?
Direct peers on record are Modal Labs, Replicate, Groq, DeepInfra, Together AI, Lepton AI, Anyscale and Fireworks AI. Hugging Face is listed as a broad incumbent. Cerebras Systems is listed as an emerging player.
Does Kog have an API?
No public API is recorded for Kog.
What industry is Kog in?
Kog's product category is AI Inference Infrastructure. Its primary akta.pro industry code is HDAAABAF, Model Deployment, Serving & Inference Platforms, with a secondary code of HDAAAAAI, AI Compiler, Runtime & Kernel Optimization Software (CUDA/ROCm/XLA, graph compilers). Its NAICS code is 5132 and its SIC code is 7372.