Developer docs
API playgroundTry for free, no card

Search company profiles

Kog

Full company profile

uuid000aqrs

Namestring
Kog
Legal namestring
Kog Labs
Websiteurl
kog.ai
Company typeenum
Private
Founded yearint
2023
Descriptiontext

Kog (legal name: Kog Labs; DBA: Kog AI) is a Paris-based AI infrastructure startup founded in 2023 by Gaël Delalleau, an École Polytechnique engineer with a background in cybersecurity research and high-performance GPU work. The company builds a hardware-software co-designed inference engine — the Kog Inference Engine (KIE) — purpose-built to eliminate software-imposed speed ceilings on standard datacenter GPUs. KIE integrates a proprietary Laneformer Transformer architecture featuring Delayed Tensor Parallelism (DTP), a single-kernel 'monokernel' runtime, the custom Kog Communication Library (KCCL), and topology-aware GPU memory access; all components are written from scratch in low-level CUDA/HIP and assembly. The system achieves 3,000 output tokens/s per request on 8× AMD MI300X and 2,100 on 8× NVIDIA H200 at batch size 1 in FP16, without speculative decoding, and is exposed as a drop-in replacement for vLLM.

Kog monetizes via API access sold to teams building AI coding agents and agentic workflows, supplemented by a high-touch Design Partner Program for enterprise engagements; no public pricing tiers are disclosed. The company also publishes the Laneformer 2B (2.3B-parameter) coding model on Hugging Face as both a usable checkpoint and a developer-acquisition channel. Distribution is API-first with a live public playground (playground.kog.ai), a technical blog (blog.kog.ai), and featured presence at AMD AI DevDay 2026. The company is backed by Varsity VC and BPI France's Deep Tech Program ($5M seed) and was awarded the French Tech 2030 label by the French government in October 2025; current operations are anchored by an 11-person team.

Short descriptiontext

Kog is a Paris-based AI infrastructure startup (founded 2023) building a hardware-software co-designed inference engine that delivers 3,000 tokens/s per request on AMD MI300X GPUs for AI coding agents and agentic workflows, via a proprietary monokernel runtime, custom KCCL communication library, and Laneformer architecture.

Operating statusenum
Operating
Ownership categoryenum
Headcount rangeband
11–50
akta.pro rankint
HeadquartersParis, France
HQ citystring
Paris
HQ countrystring
France
HQ regionstring
Europe
Markets served

Serves global market

Offices1 record

Each record includes

City, Country, Type, Description, Source

Keyword5 values
real-time LLM inference, AI inference engine, GPU optimization software, AI coding agents, low-latency inference
Industry4 codes
1Model Deployment, Serving & Inference Platforms
CodeHDAAABAFPrimaryYes
2AI Compiler, Runtime & Kernel Optimization Software (CUDA/ROCm/XLA, graph compilers)
CodeHDAAAAAIPrimaryNo
3Model Serving, Inference & Deployment Platforms (APIs, Edge/On-Prem)
CodeHDAEANACPrimaryNo
4Model Hosting, Serving & Inference Platforms
CodeHDAAACABPrimaryNo
NAICS code3 codes
  • Software Publishers5132
  • Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services51821
  • Custom Computer Programming Services541511
SIC code1 code
  • Services-Prepackaged Software7372
Product category
AI Inference Infrastructure
Social media profiles1 record
GTM motion4 records

Each record includes

Type, Description, Source

Revenue model1 record
1API Access / Inference-as-a-Service
TypeUsage Based
Description

Kog likely generates revenue by providing API access to its high-speed inference engine. The website prominently features 'Request API Access' CTAs and a public tech preview playground, suggesting a consumption-based or subscription API pricing model targeting developers and enterprises building AI agents. No specific pricing tiers are publicly disclosed.

blog.kog.ai
Marketing channels6 records

Each record includes

Title, Type, Stage, Description, Source

Distribution channels3 records

Each record includes

Title, Type, Scope, Target buyer, Description, Source

Cost components1 value
Personnel
Pricing details1 tier
1API access for developers and enterprises
ModelUsage-basedBilling cadencePay-as-you-go
Notes

Pricing details are not publicly disclosed. Access is obtained via 'Request API Access' or the Design Partner Program.

kog.ai
GTM typeB2B
B2B
Offering typeSoftware
Software
Brand1 of 4 records shown
1Kog Inference Engine (KIE)
Description

Hardware-software co-design inference engine enabling 3,000 tokens/s per request on AMD MI300X GPUs.

blog.kog.ai
+3 more records
Core offering1 text field

Kog builds a hardware-software co-designed inference engine (KIE) for ultra-low-latency LLM inference on standard datacenter GPUs (AMD MI300X and NVIDIA H200), achieving 3,000 output tokens per second per request. The platform combines a custom Laneformer model architecture with Delayed Tensor Parallelism (DTP), a monokernel GPU runtime, and a custom KCCL communication library, exposing inference through a vLLM-compatible API and a live playground. The offering targets enterprises building AI coding agents and agentic workflows where token generation speed governs agent iteration throughput.

Differentiator
Functional benefit
Problem solved
Quantifiable outcome1 of 7 values shown
  • 3,000 output tokens/s/request on AMD MI300X and 2,100 on NVIDIA H200 at batch size 1, FP16, no speculative decoding
+6 more records
Product overview1 text field

Kog is a Paris-based AI infrastructure startup offering a unified inference platform centered on the Kog Inference Engine (KIE), a hardware-software co-designed system for ultra-low-latency LLM inference. The core product integrates the Laneformer model architecture with Delayed Tensor Parallelism (DTP), the custom Kog Communication Library (KCCL), and a monokernel runtime. The platform includes Kog Labs for research/blog content and a live Playground demo. Users can access the inference engine via API request or through the drop-in vLLM-compatible interface, targeting AI coding agents and agentic workflows that require real-time token generation speeds of 3,000 tokens/s per request.

Product and service4 records
1Kog Inference Engine (KIE)
CategoryAI Inference Engine
Description

A hardware-software co-designed inference engine built to run GPUs at their absolute ceiling, delivering 3,000 tokens/s per request via a monokernel runtime, custom KCCL communication library, and Laneformer architecture. Functions as a drop-in vLLM replacement with no code refactoring required, targeting enterprise teams building AI coding agents and agentic workflows.

2Laneformer 2B
CategoryAI Model
Description

A 2.3B-parameter instruction-tuned coding model built around Kog's Delayed Tensor Parallelism architecture, designed for AI coding agents and software engineering workflows. Achieves 45.1% on HumanEval+ and 51.6% on MBPP+ while running at 3,000 tokens/s on AMD MI300X.

3Kog Laneformer Architecture
CategoryAI Architecture
Description

A novel Transformer architecture variant in which inter-device communication is delayed by one layer so that compute runs continuously without synchronization pauses, enabling linear scaling across high-end GPUs for low-latency inference.

4Kog Communication Library (KCCL)
CategoryGPU Communication Library
Description

A custom collective communication layer that replaces standard NCCL/RCCL to unlock linear scaling for tensor parallelism across high-end GPUs, achieving sub-3 microsecond AllReduce latency on AMD MI300X.

Scale indicator6 records

Each record includes

Type, Value, Description, Source

Partnership1 partner
Strategic tierCoreTypeTechnology or Integration
Description

AMD featured Kog as an ecosystem partner at AMD AI DevDay 2026 and published Kog's inference benchmark results on AMD's official engineering blog, validating Kog's 3.5x speed improvement on AMD Instinct MI300X GPUs. AMD provides the primary hardware platform for Kog's performance differentiation.

Recent move6 records

Each record includes

Date, Type, Title, Description, Source

Expansion highlight6 records

Each record includes

Type, Description

Peers10 records
TypeDirect peer
Description

Modal provides serverless GPU compute and inference infrastructure for AI applications with a developer-first, Python-native model. Comparable in targeting AI agent builders with low-friction GPU access and inference APIs.

TypeDirect peer
Description

Replicate runs a cloud API for running open-source ML models, including LLM inference, with a focus on simplicity and developer ergonomics. Comparable in API-first inference distribution to developers building AI applications.

TypeDirect peer
Description

Groq sells LPU-based ultra-low-latency inference-as-a-service, directly competing on tokens-per-second-per-request as the primary value proposition. Targets developer and enterprise agentic workloads with similar speed-first positioning.

TypeBroad incumbent
Description

Hugging Face operates Inference Endpoints and text-generation-inference (TGI), serving open-source models via API. Comparable as an inference platform serving the same open-model ecosystem, but as a broad incumbent rather than a speed-specialized player.

TypeDirect peer
Description

DeepInfra provides serverless inference APIs for open-source LLMs with custom inference engine optimizations. Directly comparable API-first inference-as-a-service business model targeting developer adoption.

TypeEmerging player
Description

Cerebras builds purpose-built inference silicon (CS-3/WSE) and a cloud inference service claiming class-leading tokens/second for LLM serving. Comparable in the speed-first inference narrative, but pursues a custom-silicon route rather than Kog's software-only approach on commodity GPUs.

TypeDirect peer
Description

Together AI runs an inference cloud with proprietary optimizations on multi-GPU stacks, including custom kernels and inference engine work. Comparable as an open-model inference API competitor targeting developer and enterprise agent builders.

TypeDirect peer
Description

Lepton AI offers a cloud platform for running AI models with custom inference optimizations on multi-GPU hardware. Comparable as a developer-facing inference platform emphasizing performance and flexible deployment.

TypeDirect peer
Description

Anyscale (Ray) offers distributed compute and serving infrastructure optimized for AI workloads, including LLM inference. Comparable as a developer-focused platform for scaling and serving large models on multi-GPU clusters.

TypeDirect peer
Description

Fireworks AI operates a model inference and fine-tuning platform optimized for low-latency serving of open-source LLMs. Directly comparable in API-first inference distribution and agent workload targeting.

Market position
Strengths5 records

Each record includes

Headline, Details, Source

Weaknesses5 records

Each record includes

Headline, Details, Source

Competitive moat5 records

Each record includes

Type, Details

Key risks6 records

Each record includes

Headline, Details, Source

Key highlights7 records

Each record includes

Headline, Details, Source

Customer concentration

Classification, Details

Segment2 records

Each record includes

Title, Type, Primary, Description, Pain point addressed, Use case, Source

Ideal customer profile2 records

Each record includes

Profile, Firmographic size, Sales motion, Sales cycle length, Buying structure, Purchase trigger, Buyer persona, Geography, Industry vertical, Primary use case, Description, Pain points, Evidence proof points, Target buyer

Technology focused
Yes
API detail
Has APIbool
No

Docs URL, Description

Integration2 records

Each record includes

Title, Type, Description, Source

AI capability5 records

Each record includes

Type, Description, Source

AI maturity
App detail

Has app

Feature8 records

Each record includes

Title, Differentiator, Description, Source

Core technology
Revenue estimate
Valuation estimate
Number of profiles
Profiles1 record

Each record includes

Name, Designation, Designation category, Overview, Profile commentary, Source

No data
No data
Funding overview

Funding stage, Last funding date, Total funding USD

Funding rounds1 record

Each record includes

Round, Amount USD, Date, Pre money valuation, Total investors, Investors, News

Investors1 record

Each record includes

Name, Type, Date of entry, Rounds participated, Website

Funding detail is available on the Subscription and Enterprise plan.Contact sales →

M&A

Each record includes

Name, Acquisition type, Announced date, Completed date, Status, Website, News

Investment

Each record includes

Name, Round, Announced date, Lead investor, Website, News

M&A and investment is available on the Subscription and Enterprise plan.Contact sales →

Kog

AI Inference Infrastructurekog.ai

Kog is a Paris-based AI infrastructure startup (founded 2023) building a hardware-software co-designed inference engine that delivers 3,000 tokens/s per request on AMD MI300X GPUs for AI coding agents and agentic workflows, via a proprietary monokernel runtime, custom KCCL communication library, and Laneformer architecture.

What Kog does

Kog (legal name: Kog Labs; DBA: Kog AI) is a Paris-based AI infrastructure startup founded in 2023 by Gaël Delalleau, an École Polytechnique engineer with a background in cybersecurity research and high-performance GPU work. The company builds a hardware-software co-designed inference engine — the Kog Inference Engine (KIE) — purpose-built to eliminate software-imposed speed ceilings on standard datacenter GPUs. KIE integrates a proprietary Laneformer Transformer architecture featuring Delayed Tensor Parallelism (DTP), a single-kernel 'monokernel' runtime, the custom Kog Communication Library (KCCL), and topology-aware GPU memory access; all components are written from scratch in low-level CUDA/HIP and assembly. The system achieves 3,000 output tokens/s per request on 8× AMD MI300X and 2,100 on 8× NVIDIA H200 at batch size 1 in FP16, without speculative decoding, and is exposed as a drop-in replacement for vLLM.

Kog monetizes via API access sold to teams building AI coding agents and agentic workflows, supplemented by a high-touch Design Partner Program for enterprise engagements; no public pricing tiers are disclosed. The company also publishes the Laneformer 2B (2.3B-parameter) coding model on Hugging Face as both a usable checkpoint and a developer-acquisition channel. Distribution is API-first with a live public playground (playground.kog.ai), a technical blog (blog.kog.ai), and featured presence at AMD AI DevDay 2026. The company is backed by Varsity VC and BPI France's Deep Tech Program ($5M seed) and was awarded the French Tech 2030 label by the French government in October 2025; current operations are anchored by an 11-person team.

Kog firmographics

Firmographics
Name
Kog
Legal name
Kog Labs
Website
https://kog.ai
Company type
Private
Founded year
2023
Operating status
Operating
Headcount range
11–50 employees
Short description
Kog is a Paris-based AI infrastructure startup (founded 2023) building a hardware-software co-designed inference engine that delivers 3,000 tokens/s per request on AMD MI300X GPUs for AI coding agents and agentic workflows, via a proprietary monokernel runtime, custom KCCL communication library, and Laneformer architecture.
Ownership category
akta.pro rank

Kog industry classification

Industry
Product category
AI Inference Infrastructure
NAICS
Software Publishers (5132), Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (51821), Custom Computer Programming Services (541511)
SIC
Services-Prepackaged Software (7372)
akta.pro primary industry
Model Deployment, Serving & Inference Platforms (HDAAABAF)
akta.pro secondary industries
AI Compiler, Runtime & Kernel Optimization Software (CUDA/ROCm/XLA, graph compilers) (HDAAAAAI), Model Serving, Inference & Deployment Platforms (APIs, Edge/On-Prem) (HDAEANAC), Model Hosting, Serving & Inference Platforms (HDAAACAB)

Keywords

  • Real-time LLM inference
  • AI inference engine
  • GPU optimization software
  • AI coding agents
  • Low-latency inference

Where Kog is headquartered

Location

Headquarters

HQ city
Paris
HQ country
France
HQ region
Europe

Offices1 record

Markets served

Kog business model

Business model
GTM type
B2B
Offering type
Software
Cost components
Personnel

Revenue model

  1. API Access / Inference-as-a-Service: Kog likely generates revenue by providing API access to its high-speed inference engine. The website prominently features 'Request API Access' CTAs and a public tech preview playground, suggesting a consumption-based or subscription API pricing model targeting developers and enterprises building AI agents. No specific pricing tiers are publicly disclosed.

Pricing tiers

ModelBillingPrice
Usage-basedPay-as-you-goAPI access for developers and enterprises

Go-to-market motion4 records

Distribution channels3 records

Marketing channels6 records

Kog product offering

Product offering

Core offering

Kog builds a hardware-software co-designed inference engine (KIE) for ultra-low-latency LLM inference on standard datacenter GPUs (AMD MI300X and NVIDIA H200), achieving 3,000 output tokens per second per request. The platform combines a custom Laneformer model architecture with Delayed Tensor Parallelism (DTP), a monokernel GPU runtime, and a custom KCCL communication library, exposing inference through a vLLM-compatible API and a live playground. The offering targets enterprises building AI coding agents and agentic workflows where token generation speed governs agent iteration throughput.

Product overview

Kog is a Paris-based AI infrastructure startup offering a unified inference platform centered on the Kog Inference Engine (KIE), a hardware-software co-designed system for ultra-low-latency LLM inference. The core product integrates the Laneformer model architecture with Delayed Tensor Parallelism (DTP), the custom Kog Communication Library (KCCL), and a monokernel runtime. The platform includes Kog Labs for research/blog content and a live Playground demo. Users can access the inference engine via API request or through the drop-in vLLM-compatible interface, targeting AI coding agents and agentic workflows that require real-time token generation speeds of 3,000 tokens/s per request.

Differentiator

Problem solved

Functional benefit

Brands

  • Kog Inference Engine (KIE): Hardware-software co-design inference engine enabling 3,000 tokens/s per request on AMD MI300X GPUs.
  • Laneformer
  • Kog CommunicationLibrary (KCCL)
  • Kog Playground

Products and services

  • Kog Inference Engine (KIE) A hardware-software co-designed inference engine built to run GPUs at their absolute ceiling, delivering 3,000 tokens/s per request via a monokernel runtime, custom KCCL communication library, and Laneformer architecture. Functions as a drop-in vLLM replacement with no code refactoring required, targeting enterprise teams building AI coding agents and agentic workflows.
  • Laneformer 2B A 2.3B-parameter instruction-tuned coding model built around Kog's Delayed Tensor Parallelism architecture, designed for AI coding agents and software engineering workflows. Achieves 45.1% on HumanEval+ and 51.6% on MBPP+ while running at 3,000 tokens/s on AMD MI300X.
  • Kog Laneformer Architecture A novel Transformer architecture variant in which inter-device communication is delayed by one layer so that compute runs continuously without synchronization pauses, enabling linear scaling across high-end GPUs for low-latency inference.
  • Kog Communication Library (KCCL) A custom collective communication layer that replaces standard NCCL/RCCL to unlock linear scaling for tensor parallelism across high-end GPUs, achieving sub-3 microsecond AllReduce latency on AMD MI300X.

Quantifiable outcome

  • 3,000 output tokens/s/request on AMD MI300X and 2,100 on NVIDIA H200 at batch size 1, FP16, no speculative decoding
  • +6 more outcomes

Companies that use Kog

Customer profile

Segments2 records

Ideal customer profiles2 records

Kog technology and API

Technology

Technology focussed Yes

API detail

Has API
No
API docs
API detail

Core technology

AI maturity

App detail

Integration2 records

AI capability5 records

Feature8 records

Kog partnerships and signals

Strategic signal

Partnerships

One partnership is on record.

  • AMDcoreTechnology or IntegrationAMD featured Kog as an ecosystem partner at AMD AI DevDay 2026 and published Kog's inference benchmark results on AMD's official engineering blog, validating Kog's 3.5x speed improvement on AMD Instinct MI300X GPUs. AMD provides the primary hardware platform for Kog's performance differentiation.

Scale indicators6 records

Recent moves6 records

Expansion highlights6 records

Kog competitors and assessment

Company assessment

Direct peers

  • Modal Labs: Modal provides serverless GPU compute and inference infrastructure for AI applications with a developer-first, Python-native model. Comparable in targeting AI agent builders with low-friction GPU access and inference APIs.
  • Replicate: Replicate runs a cloud API for running open-source ML models, including LLM inference, with a focus on simplicity and developer ergonomics. Comparable in API-first inference distribution to developers building AI applications.
  • Groq: Groq sells LPU-based ultra-low-latency inference-as-a-service, directly competing on tokens-per-second-per-request as the primary value proposition. Targets developer and enterprise agentic workloads with similar speed-first positioning.
  • DeepInfra: DeepInfra provides serverless inference APIs for open-source LLMs with custom inference engine optimizations. Directly comparable API-first inference-as-a-service business model targeting developer adoption.
  • Together AI: Together AI runs an inference cloud with proprietary optimizations on multi-GPU stacks, including custom kernels and inference engine work. Comparable as an open-model inference API competitor targeting developer and enterprise agent builders.
  • Lepton AI: Lepton AI offers a cloud platform for running AI models with custom inference optimizations on multi-GPU hardware. Comparable as a developer-facing inference platform emphasizing performance and flexible deployment.
  • Anyscale: Anyscale (Ray) offers distributed compute and serving infrastructure optimized for AI workloads, including LLM inference. Comparable as a developer-focused platform for scaling and serving large models on multi-GPU clusters.
  • Fireworks AI: Fireworks AI operates a model inference and fine-tuning platform optimized for low-latency serving of open-source LLMs. Directly comparable in API-first inference distribution and agent workload targeting.

Broad incumbents

  • Hugging Face: Hugging Face operates Inference Endpoints and text-generation-inference (TGI), serving open-source models via API. Comparable as an inference platform serving the same open-model ecosystem, but as a broad incumbent rather than a speed-specialized player.

Emerging players

  • Cerebras Systems: Cerebras builds purpose-built inference silicon (CS-3/WSE) and a cloud inference service claiming class-leading tokens/second for LLM serving. Comparable in the speed-first inference narrative, but pursues a custom-silicon route rather than Kog's software-only approach on commodity GPUs.

Market position

Strengths5 records

Weaknesses5 records

Competitive moat5 records

Key risks6 records

Key highlights7 records

Customer concentration

Kog social profiles

Digital presence

Kog financial estimates

Financial estimate

Revenue estimate

Valuation estimate

Kog leadership team

Management profile

Number of profiles

Profiles1 record

Kog funding detail

Funding detail

Funding overview

Funding rounds1 record

Investors1 record

Funding detail is available on the Subscription and Enterprise plan.Contact sales →

Kog M&A and investment

M&A and investment

M&A

Investments

M&A and investment is available on the Subscription and Enterprise plan.Contact sales →

Frequently asked questions about Kog

What does Kog do?

Kog builds a hardware-software co-designed inference engine (KIE) for ultra-low-latency LLM inference on standard datacenter GPUs (AMD MI300X and NVIDIA H200), achieving 3,000 output tokens per second per request. The platform combines a custom Laneformer model architecture with Delayed Tensor Parallelism (DTP), a monokernel GPU runtime, and a custom KCCL communication library, exposing inference through a vLLM-compatible API and a live playground. The offering targets enterprises building AI coding agents and agentic workflows where token generation speed governs agent iteration throughput.

Is Kog a public or private company?

Kog is a private company. It is classified as venture growth investor backed and is currently operating.

When was Kog founded?

Kog was founded in 2023. It employs 11 to 50 people.

Where is Kog based?

Kog is headquartered in Paris, France, in the Europe region.

How does Kog make money?

One revenue line is on record: API Access / Inference-as-a-Service.

Who are Kog's main competitors?

Direct peers on record are Modal Labs, Replicate, Groq, DeepInfra, Together AI, Lepton AI, Anyscale and Fireworks AI. Hugging Face is listed as a broad incumbent. Cerebras Systems is listed as an emerging player.

Does Kog have an API?

No public API is recorded for Kog.

What industry is Kog in?

Kog's product category is AI Inference Infrastructure. Its primary akta.pro industry code is HDAAABAF, Model Deployment, Serving & Inference Platforms, with a secondary code of HDAAAAAI, AI Compiler, Runtime & Kernel Optimization Software (CUDA/ROCm/XLA, graph compilers). Its NAICS code is 5132 and its SIC code is 7372.

Unlock the full company data

50 free credits on sign-up, no credit card required.

Contact sales
Live signals
TechbuzzFrench Startup Kog Challenges GPU Inference LimitsFrench startup Kog is challenging the industry consensus that GPUs are inefficient for agentic AI workflows by optimizing software orchestration to improve utilization on existing hardware. The company aims to address rising infrastructure costs and GPU shortages by enabling better multitasking of multi-step AI processes without requiring new silicon. This strategy positions Kog against competitors developing specialized chips and major cloud providers building custom AI hardware.BEAMSTART: NewsKog Startup Aims to Deliver Lightning-Fast AI on Everyday GPUsKog, a French startup, is developing software optimizations to significantly increase the speed of large language model inference on standard NVIDIA and AMD GPUs. The company demonstrated performance of 3,000 tokens per second on a small model and aims to achieve ten times faster speeds on larger models by September. This approach seeks to reduce reliance on expensive specialized hardware and lower energy consumption in the AI sector.TechCrunchKog is going deeper to squeeze more inference out of GPUsFrench startup Kog announced its Kog Inference Engine (KIE), a software optimization tool designed to accelerate Large Language Model inference on standard datacenter GPUs like those from AMD and Nvidia. The company aims to achieve significant speed improvements, such as 30x faster processing, to address current bottlenecks in AI workflows for enterprise customers. CEO Gaël Delalleau plans to demonstrate traction with major models by September to secure Series A funding.HuggingfaceKog Laneformer 2B: The Latency-First Model Behind Kog Inference EngineKog, a Paris-based AI infrastructure startup, has released the weights and model code of Laneformer 2B, a 2.3B-parameter instruction-tuned coding model, on Hugging Face Hub as both a usable checkpoint and research artifact. The model was designed from the ground up for latency optimization, featuring a lane-structured Transformer architecture with Delayed Tensor Parallelism (DTP) and reaching 45.1% HumanEval+ and 51.6% MBPP+ benchmarks while achieving 3,000 output tokens/s/request on AMD MI300X and 2,100 output tokens/s/request on NVIDIA H200 GPUs. The training utilized 192 NVIDIA H100 GPUs across Scaleway and ADASTRA infrastructure over approximately 21 days, consuming about 6T tokens total for pre-training, mid-training, and post-training phases.The French Tech JournalThe AI Industry Spent Billions Chasing Faster Chips. Inference Startup Kog Says They Were Solving the Wrong Problem.Kog, a Paris-based AI infrastructure startup, opened a public tech preview of its inference engine, claiming to achieve dedicated-silicon speeds on standard GPUs. On a single node of eight AMD MI300X GPUs, it generates over 3,000 output tokens per second per user request. The company argues that better software may be sufficient to avoid migrating to new hardware.