Developer docs
API playgroundTry for free, no card

Search company profiles

Infinigence

Full company profile

uuid0000u9h

Namestring
Infinigence
Legal namestring
Infinigence
Company typeenum
Private
Founded yearint
2023
Descriptiontext

Infinigence (无问芯穹) is a Beijing-based AI infrastructure company founded in 2023, affiliated with Tsinghua University's Department of Electronic Engineering. It provides AGI computing power solutions centered on large model energy efficiency optimization. Its core product, the Agentic MaaS Platform, operates as a neutral infrastructure layer between chip manufacturers and AI model developers, enabling AI inference workloads with prefill-decode separation technology that delivers a stated 5-10x cost-performance improvement. The platform also facilitates deployment of AI models on domestic Chinese chips, addressing hardware compatibility constraints in the Chinese AI ecosystem.

The company runs an API-first, token-based revenue model, generating income from token consumption on its inference platform. Its primary customer segments are enterprise AI agent operators and AI model developers seeking scalable, cost-efficient inference infrastructure. The company is headquartered in Beijing with 51-100 employees and has raised approximately $244 million across four funding rounds since December 2023, with backers including Legend Capital, Qiming Venture Partners, Shunwei Capital, Xiaomi, Shanghai Guotou, and Alibaba Entrepreneurs Fund.

Token call volume grew over 20x from December 2025 to April 2026 and doubled roughly every two weeks since late January 2026, indicating rapid adoption of the platform's inference services. Infinigence operates exclusively in China, with no disclosed international expansion.

Short descriptiontext

Infinigence (无问芯穹) is a Beijing-based AI infrastructure company founded in 2023 that provides an Agentic MaaS Platform using prefill-decode separation to deliver cost-efficient inference for AI model developers and AI agent operators in China.

Operating statusenum
Operating
Ownership categoryenum
Headcount rangeband
51–100
akta.pro rankint
HeadquartersBeijing, China
HQ citystring
Beijing
HQ countrystring
China
HQ regionstring
Asia
Markets served

Serves global market

Offices1 record

Each record includes

City, Country, Type, Description, Source

Keyword5 values
AI inference infrastructure, model-as-a-service, GPU compute optimization, token-based API, heterogeneous chip deployment
Industry4 codes
1Model Serving, Inference & Deployment Platforms (APIs, Edge/On-Prem)
CodeHDAEANACPrimaryYes
2Model Deployment, Serving & Inference Platforms
CodeHDAAABAFPrimaryNo
3End-to-End Enterprise AI Platforms (MLOps & Model Lifecycle Management)
CodeHDAEANAAPrimaryNo
4AI Compiler, Runtime & Kernel Optimization Software (CUDA/ROCm/XLA, graph compilers)
CodeHDAAAAAIPrimaryNo
NAICS code2 codes
  • Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services51821
  • Computer Systems Design and Related Services5415
SIC code2 codes
  • Services-Prepackaged Software7372
  • Services-Computer Integrated Systems Design7373
Product category
AI Inference Infrastructure
Social media profiles1 record
GTM motion2 records

Each record includes

Type, Description, Source

Revenue model1 record
1AI Inference Compute (Token-Based)
TypeUsage Based
Description

Revenue is generated through token-based API usage on the Agentic MaaS platform. The company acts as a neutral 'token factory', providing AI inference compute services to model developers and enterprise customers. Token consumption has shown explosive growth with 20x volume increase over four months, and the broader market saw 15x token revenue growth for major providers since early 2026. The model is usage-based with consumption scaling with enterprise AI agent adoption.

pandaily.com
Marketing channels1 record

Each record includes

Title, Type, Stage, Description, Source

Distribution channels1 record

Each record includes

Title, Type, Scope, Target buyer, Description, Source

Cost components5 values
Technology or R&D, Infrastructure, Personnel, Operations, Marketing or Sales
GTM typeB2B
B2B
Offering typeSoftware
Software
Core offering1 text field

Infinigence operates an Agentic MaaS (Model-as-a-Service) platform that provides scalable AI inference compute infrastructure for AI agents and enterprise workloads. The platform functions as a neutral 'token factory' middleware between chip manufacturers and model developers, leveraging prefill-decode separation technology to achieve 5-10x cost-performance improvements while enabling AI model deployment on heterogeneous hardware, including domestic Chinese chips.

Differentiator
Functional benefit
Problem solved
Quantifiable outcome1 of 3 values shown
  • 20x token call volume growth from December 2025 to April 2026
+2 more records
Product overview1 text field

Infinigence is an AI infrastructure company affiliated with Tsinghua University's Department of Electronic Engineering, offering a single core product—the Agentic MaaS Platform—that positions itself as a neutral infrastructure layer between chip manufacturers and AI model developers. The platform facilitates AI inference compute workloads and token generation services, leveraging prefill-decode separation technology to deliver 5-10x cost-performance improvements while providing a deployment pathway for domestic Chinese chips.

Product and service1 record
1Agentic MaaS Platform
CategoryAI Inference Infrastructure
Scale indicator5 records

Each record includes

Type, Value, Description, Source

Recent move5 records

Each record includes

Date, Type, Title, Description, Source

Expansion highlight4 records

Each record includes

Type, Description

Peers10 records
TypeDirect peer
Description

Replicate provides a cloud API for running and deploying machine learning models, emphasizing simple API-based access to inference infrastructure. It serves the same developer- and enterprise-customer base seeking scalable, on-demand AI inference as Infinigence's Agentic MaaS platform.

TypeDirect peer
Description

Anyscale provides AI infrastructure built on Ray for distributed model training and serving, targeting enterprise AI workloads with cost and performance optimization. Its serving and inference capabilities directly overlap with Infinigence's positioning as an enterprise AI compute platform.

TypeDirect peer
Description

SiliconFlow is a Chinese AI inference infrastructure startup offering model serving and inference optimization across heterogeneous hardware including domestic Chinese chips. It competes directly with Infinigence's Agentic MaaS platform in the same 'token factory' positioning for Chinese AI model developers and enterprise customers.

TypeDirect peer
Description

Fireworks AI is a global AI inference platform specializing in fast, cost-efficient model serving and deployment, with proprietary optimization techniques similar to Infinigence's prefill-decode separation. It serves the same customer base of model developers and enterprise AI builders seeking production-grade inference infrastructure.

TypeDirect peer
Description

Modal Labs provides serverless cloud infrastructure for AI inference and compute workloads, with developer-friendly APIs and a focus on scalable model serving. Its API-first, usage-based model mirrors Infinigence's Agentic MaaS approach for AI model deployment and inference.

TypeBroad incumbent
Description

Baidu AI Cloud offers the Qianfan model platform and broader AI infrastructure including inference services to enterprise and developer customers in China. It is a broad incumbent competing with Infinigence for the same Chinese AI inference workloads, with the added dimension of Baidu being an investor in the company.

TypeDirect peer
Description

Together AI operates a cloud platform for training, fine-tuning, and serving open-source AI models, with significant focus on inference optimization and cost-efficient deployment. It overlaps with Infinigence in providing scalable inference infrastructure to AI developers and enterprises, though serving global rather than Chinese markets.

TypeBroad incumbent
Description

Tencent Cloud provides AI infrastructure services including model serving and inference for enterprise customers in China. As a broad hyperscale incumbent, it competes with Infinigence for enterprise AI agent workloads while offering a much wider portfolio of services.

TypeBroad incumbent
Description

Alibaba Cloud is a hyperscale cloud provider offering Model Studio (Bailian) and PAI for AI training and inference at massive scale. While serving a broader portfolio, its inference capabilities compete directly with Infinigence for Chinese enterprise AI customers and the same underlying compute workloads.

TypeBroad incumbent
Description

Volcano Engine is ByteDance's cloud and AI platform offering model serving, inference infrastructure, and AI development services to enterprise customers. It competes with Infinigence in the Chinese enterprise AI infrastructure market while leveraging ByteDance's internal AI scale and capabilities.

Market position
Strengths5 records

Each record includes

Headline, Details, Source

Weaknesses5 records

Each record includes

Headline, Details, Source

Competitive moat4 records

Each record includes

Type, Details

Key risks6 records

Each record includes

Headline, Details, Source

Key highlights7 records

Each record includes

Headline, Details, Source

Customer concentration

Classification, Details

Segment2 records

Each record includes

Title, Type, Primary, Description, Pain point addressed, Use case, Source

Ideal customer profile2 records

Each record includes

Profile, Firmographic size, Sales motion, Sales cycle length, Buying structure, Purchase trigger, Buyer persona, Geography, Industry vertical, Primary use case, Description, Pain points, Evidence proof points, Target buyer

Technology focused
Yes
API detail
Has APIbool
No

Docs URL, Description

AI maturity
App detail

Has app

Feature3 records

Each record includes

Title, Differentiator, Description, Source

Core technology
Revenue estimate
Valuation estimate
Number of profiles
No data
No data
Funding overview

Funding stage, Last funding date, Total funding USD

Funding rounds4 records

Each record includes

Round, Amount USD, Date, Pre money valuation, Total investors, Investors, News

Investors43 records

Each record includes

Name, Type, Date of entry, Rounds participated, Website

Funding detail is available on the Subscription and Enterprise plan.Contact sales →

M&A

Each record includes

Name, Acquisition type, Announced date, Completed date, Status, Website, News

Investment

Each record includes

Name, Round, Announced date, Lead investor, Website, News

M&A and investment is available on the Subscription and Enterprise plan.Contact sales →

Infinigence

AI Inference Infrastructurecloud.infini-ai.com

Infinigence (无问芯穹) is a Beijing-based AI infrastructure company founded in 2023 that provides an Agentic MaaS Platform using prefill-decode separation to deliver cost-efficient inference for AI model developers and AI agent operators in China.

What Infinigence does

Infinigence (无问芯穹) is a Beijing-based AI infrastructure company founded in 2023, affiliated with Tsinghua University's Department of Electronic Engineering. It provides AGI computing power solutions centered on large model energy efficiency optimization. Its core product, the Agentic MaaS Platform, operates as a neutral infrastructure layer between chip manufacturers and AI model developers, enabling AI inference workloads with prefill-decode separation technology that delivers a stated 5-10x cost-performance improvement. The platform also facilitates deployment of AI models on domestic Chinese chips, addressing hardware compatibility constraints in the Chinese AI ecosystem.

The company runs an API-first, token-based revenue model, generating income from token consumption on its inference platform. Its primary customer segments are enterprise AI agent operators and AI model developers seeking scalable, cost-efficient inference infrastructure. The company is headquartered in Beijing with 51-100 employees and has raised approximately $244 million across four funding rounds since December 2023, with backers including Legend Capital, Qiming Venture Partners, Shunwei Capital, Xiaomi, Shanghai Guotou, and Alibaba Entrepreneurs Fund.

Token call volume grew over 20x from December 2025 to April 2026 and doubled roughly every two weeks since late January 2026, indicating rapid adoption of the platform's inference services. Infinigence operates exclusively in China, with no disclosed international expansion.

Infinigence firmographics

Firmographics
Name
Infinigence
Legal name
Infinigence
Website
https://cloud.infini-ai.com
Company type
Private
Founded year
2023
Operating status
Operating
Headcount range
51–100 employees
Short description
Infinigence (无问芯穹) is a Beijing-based AI infrastructure company founded in 2023 that provides an Agentic MaaS Platform using prefill-decode separation to deliver cost-efficient inference for AI model developers and AI agent operators in China.
Ownership category
akta.pro rank

Infinigence industry classification

Industry
Product category
AI Inference Infrastructure
NAICS
Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (51821), Computer Systems Design and Related Services (5415)
SIC
Services-Prepackaged Software (7372), Services-Computer Integrated Systems Design (7373)
akta.pro primary industry
Model Serving, Inference & Deployment Platforms (APIs, Edge/On-Prem) (HDAEANAC)
akta.pro secondary industries
Model Deployment, Serving & Inference Platforms (HDAAABAF), End-to-End Enterprise AI Platforms (MLOps & Model Lifecycle Management) (HDAEANAA), AI Compiler, Runtime & Kernel Optimization Software (CUDA/ROCm/XLA, graph compilers) (HDAAAAAI)

Keywords

  • AI inference infrastructure
  • Model-as-a-service
  • GPU compute optimization
  • Token-based API
  • Heterogeneous chip deployment

Where Infinigence is headquartered

Location

Headquarters

HQ city
Beijing
HQ country
China
HQ region
Asia

Offices1 record

Markets served

Infinigence business model

Business model
GTM type
B2B
Offering type
Software
Cost components
Technology or R&D, Infrastructure, Personnel, Operations, Marketing or Sales

Revenue model

  1. AI Inference Compute (Token-Based): Revenue is generated through token-based API usage on the Agentic MaaS platform. The company acts as a neutral 'token factory', providing AI inference compute services to model developers and enterprise customers. Token consumption has shown explosive growth with 20x volume increase over four months, and the broader market saw 15x token revenue growth for major providers since early 2026. The model is usage-based with consumption scaling with enterprise AI agent adoption.

Go-to-market motion2 records

Distribution channels1 record

Marketing channels1 record

Infinigence product offering

Product offering

Core offering

Infinigence operates an Agentic MaaS (Model-as-a-Service) platform that provides scalable AI inference compute infrastructure for AI agents and enterprise workloads. The platform functions as a neutral 'token factory' middleware between chip manufacturers and model developers, leveraging prefill-decode separation technology to achieve 5-10x cost-performance improvements while enabling AI model deployment on heterogeneous hardware, including domestic Chinese chips.

Product overview

Infinigence is an AI infrastructure company affiliated with Tsinghua University's Department of Electronic Engineering, offering a single core product—the Agentic MaaS Platform—that positions itself as a neutral infrastructure layer between chip manufacturers and AI model developers. The platform facilitates AI inference compute workloads and token generation services, leveraging prefill-decode separation technology to deliver 5-10x cost-performance improvements while providing a deployment pathway for domestic Chinese chips.

Differentiator

Problem solved

Functional benefit

Products and services

  • Agentic MaaS Platform

Quantifiable outcome

  • 20x token call volume growth from December 2025 to April 2026
  • +2 more outcomes

Companies that use Infinigence

Customer profile

Segments2 records

Ideal customer profiles2 records

Infinigence technology and API

Technology

Technology focussed Yes

API detail

Has API
No
API docs
API detail

Core technology

AI maturity

App detail

Feature3 records

Infinigence partnerships and signals

Strategic signal

Scale indicators5 records

Recent moves5 records

Expansion highlights4 records

Infinigence competitors and assessment

Company assessment

Direct peers

  • Replicate: Replicate provides a cloud API for running and deploying machine learning models, emphasizing simple API-based access to inference infrastructure. It serves the same developer- and enterprise-customer base seeking scalable, on-demand AI inference as Infinigence's Agentic MaaS platform.
  • Anyscale: Anyscale provides AI infrastructure built on Ray for distributed model training and serving, targeting enterprise AI workloads with cost and performance optimization. Its serving and inference capabilities directly overlap with Infinigence's positioning as an enterprise AI compute platform.
  • SiliconFlow (硅基流动): SiliconFlow is a Chinese AI inference infrastructure startup offering model serving and inference optimization across heterogeneous hardware including domestic Chinese chips. It competes directly with Infinigence's Agentic MaaS platform in the same 'token factory' positioning for Chinese AI model developers and enterprise customers.
  • Fireworks AI: Fireworks AI is a global AI inference platform specializing in fast, cost-efficient model serving and deployment, with proprietary optimization techniques similar to Infinigence's prefill-decode separation. It serves the same customer base of model developers and enterprise AI builders seeking production-grade inference infrastructure.
  • Modal Labs: Modal Labs provides serverless cloud infrastructure for AI inference and compute workloads, with developer-friendly APIs and a focus on scalable model serving. Its API-first, usage-based model mirrors Infinigence's Agentic MaaS approach for AI model deployment and inference.
  • Together AI: Together AI operates a cloud platform for training, fine-tuning, and serving open-source AI models, with significant focus on inference optimization and cost-efficient deployment. It overlaps with Infinigence in providing scalable inference infrastructure to AI developers and enterprises, though serving global rather than Chinese markets.

Broad incumbents

  • Baidu AI Cloud (Qianfan): Baidu AI Cloud offers the Qianfan model platform and broader AI infrastructure including inference services to enterprise and developer customers in China. It is a broad incumbent competing with Infinigence for the same Chinese AI inference workloads, with the added dimension of Baidu being an investor in the company.
  • Tencent Cloud: Tencent Cloud provides AI infrastructure services including model serving and inference for enterprise customers in China. As a broad hyperscale incumbent, it competes with Infinigence for enterprise AI agent workloads while offering a much wider portfolio of services.
  • Alibaba Cloud (Model Studio / PAI): Alibaba Cloud is a hyperscale cloud provider offering Model Studio (Bailian) and PAI for AI training and inference at massive scale. While serving a broader portfolio, its inference capabilities compete directly with Infinigence for Chinese enterprise AI customers and the same underlying compute workloads.
  • Volcano Engine (ByteDance): Volcano Engine is ByteDance's cloud and AI platform offering model serving, inference infrastructure, and AI development services to enterprise customers. It competes with Infinigence in the Chinese enterprise AI infrastructure market while leveraging ByteDance's internal AI scale and capabilities.

Market position

Strengths5 records

Weaknesses5 records

Competitive moat4 records

Key risks6 records

Key highlights7 records

Customer concentration

Infinigence social profiles

Digital presence

Infinigence financial estimates

Financial estimate

Revenue estimate

Valuation estimate

Infinigence leadership team

Management profile

Number of profiles

Infinigence funding detail

Funding detail

Funding overview

Funding rounds4 records

Investors43 records

Funding detail is available on the Subscription and Enterprise plan.Contact sales →

Infinigence M&A and investment

M&A and investment

M&A

Investments

M&A and investment is available on the Subscription and Enterprise plan.Contact sales →

Frequently asked questions about Infinigence

What does Infinigence do?

Infinigence operates an Agentic MaaS (Model-as-a-Service) platform that provides scalable AI inference compute infrastructure for AI agents and enterprise workloads. The platform functions as a neutral 'token factory' middleware between chip manufacturers and model developers, leveraging prefill-decode separation technology to achieve 5-10x cost-performance improvements while enabling AI model deployment on heterogeneous hardware, including domestic Chinese chips.

Is Infinigence a public or private company?

Infinigence is a private company. It is classified as venture growth investor backed and is currently operating.

When was Infinigence founded?

Infinigence was founded in 2023. It employs 51 to 100 people.

Where is Infinigence based?

Infinigence is headquartered in Beijing, China, in the Asia region.

How does Infinigence make money?

One revenue line is on record: AI Inference Compute (Token-Based).

Who are Infinigence's main competitors?

Direct peers on record are Replicate, Anyscale, SiliconFlow (硅基流动), Fireworks AI, Modal Labs and Together AI. Broad incumbents are Baidu AI Cloud (Qianfan), Tencent Cloud, Alibaba Cloud (Model Studio / PAI) and Volcano Engine (ByteDance).

Does Infinigence have an API?

No public API is recorded for Infinigence.

What industry is Infinigence in?

Infinigence's product category is AI Inference Infrastructure. Its primary akta.pro industry code is HDAEANAC, Model Serving, Inference & Deployment Platforms (APIs, Edge/On-Prem), with a secondary code of HDAAABAF, Model Deployment, Serving & Inference Platforms. Its NAICS code is 51821 and its SIC code is 7372.

Unlock the full company data

50 free credits on sign-up, no credit card required.

Contact sales
Live signals
Tech in AsiaTencent-backed Infinigence eyes HK IPO as early as 2027Shanghai-based Infinigence AI is preparing for a Hong Kong IPO as early as the first half of 2027 after a confidential filing to raise several hundred million US dollars. The company has raised 4.3 billion yuan (US$642 million) since 2023 and was valued at 14.3 billion yuan (US$2.13 billion) ahead of the offering.Crypto BriefingInfinigence AI eyes Hong Kong IPO as China’s demand for AI compute growsInfinigence AI, a Tsinghua-linked cloud infrastructure startup backed by Tencent and Baidu, plans a Hong Kong IPO as early as the first half of 2027. The company has raised about $641 million and operates 53 data centers across 26 cities. It aims to supply heterogeneous computing infrastructure for China's growing AI demand.Business News TodayChina’s Infinigence AI plans Hong Kong IPO to tap surging AI computing demand- Moneycontrol.comInfinigence AI, a Chinese AI cloud infrastructure provider, is preparing for a Hong Kong IPO in the first half of next year. The company has raised 4.3 billion yuan ($641 million) since 2023 and is valued at 14.3 billion yuan. Details on timing and size are still being finalized.BloombergTencent-Backed AI Cloud Unicorn Said to File for Hong Kong IPOInfinigence AI, a Chinese AI cloud infrastructure provider, is preparing for a Hong Kong IPO in the first half of next year. The company has raised 4.3 billion yuan since 2023 and is valued at 14.3 billion yuan. It plans to raise several hundred million US dollars.PandailyInfinigence AI Open-Sources APXInf for Embodied Edge Inference on Jetson ThorInfinigence AI open-sourced APXInf, an edge inference engine for embodied models, with a Rust runtime and Python bindings for Jetson-class devices. On NVIDIA Jetson Thor, FP8 inference drops from about 278 ms to under 26 ms, roughly 38.46 Hz, a 10× latency reduction. The release targets Jetson Orin, Thor, and GeForce RTX 4090, with a roadmap for more model ports and AMD backends.LeiphoneThe Primitive Rhythm and Wuwen Core have reached a strategic cooperation to promote high-quality Token supply and application.TokenRhythm and Infinigence AI signed a strategic cooperation agreement on September 15, 2026, to combine Infinigence AI's Agentic Infra with TokenRhythm's multi-model services and intelligent routing. The partnership aims to supply, schedule, and apply high-quality tokens, with joint verification in key industries and enterprise-level Agent delivery. Both parties plan to extend cooperation to enterprise solution design and customer expansion.LeiphoneNo Questions About the Core: When Enterprises Encounter Token 'Assassins', We Decide to 'Open Stores and Build Centers' | WAIC 2026Wuwen Chip, an AI infrastructure firm, announced at WAIC 2026 its 'front store, back factory, one center' Agentic Infra strategy to address 'token assassins' and optimize heterogeneous chips. The company reported a tenfold reduction in inference costs and a 40-fold increase in daily token call volume. It also laid out embodied intelligence as a future direction.Leiphone不止 Token 工厂,无问芯穹“前店后厂一中心”Agentic Infra 战略布局重磅发布Wumen Xinqiong unveiled its Agentic Infra platform at the 2026 World AI Conference, covering cross-cluster training and inference optimizations. The system achieved 165% performance improvement in ultra-long-distance training and 51.5% reduction in first-token latency. It plans to support over 100,000 cards for next-generation Agentic AI.Pandaily20x Token Growth in Six Months: Inside a Chinese 'Token Factory's' Business FlywheelInfinigence, a Chinese AI infrastructure company affiliated with Tsinghua University's Department of Electronic Engineering, reported that its Agentic MaaS platform achieved token call volume growth exceeding 20x from December to April, driven by a structural shift where inference has overtaken training as the dominant AI compute workload. The company positions itself as a neutral 'token factory' between chip manufacturers and model developers, achieving 5-10x cost-performance improvements through prefill-decode separation technology, which also creates a deployment pathway for domestic Chinese chips. Global enterprise spending on inference infrastructure is projected to reach $68 billion in 2026, compared to $45 billion for training infrastructure.South China Morning PostChina leads US in everyday AI apps but firms are overvalued, experts sayChinese labs have unveiled new models trailing US peers like OpenAI and Anthropic, but China leads the US in everyday AI apps. Alibaba Cloud's Chi Zhang said China is only 100 days behind in frontier AI capabilities, though firms are increasingly overvalued.