FlexAI
FlexAI is a Paris-based AI infrastructure company offering an OpenAI-compatible platform for running 30+ open-weight models on heterogeneous NVIDIA/AMD GPUs across multiple clouds, with serverless, dedicated, and private-cloud deployment tiers for builders, growing AI teams, and regulated enterprises.
- Company typePrivate
- Founded2023
- HeadquartersParis, France
- Headcount11–50
- GTM typeB2B
- OfferingSoftware
What FlexAI does
FlexAI is a Paris-based AI infrastructure company founded in 2023 by ex-NVIDIA engineers and led by CEO Brijesh Tripathi, who previously deployed the Aurora supercomputer and managed 50,000+ GPUs at Intel. The company sells an "agent-native" compute platform that exposes 30+ open-weight models (Qwen, DeepSeek, Llama, Mistral, Gemma, GPT-OSS, Nemotron, GLM) through a single OpenAI-compatible API across text, vision, image, audio, and embedding modalities, with hardware-aware scheduling over heterogeneous fleets of NVIDIA (H100/H200/B200) and AMD GPUs running across AWS, Google Cloud, and on-premises environments.
The product stack is organized in three tiers. Token Factory is the serverless, pay-per-token inference tier accessed via tokens.flex.ai. Dedicated Endpoints reserve H100/H200 capacity for production with autoscaling and scale-to-zero on per-second billing (marketed at $2.10/hr for H100 versus $2.50–$6.98/hr at hyperscalers). AI Factory delivers the same capability inside a VPC, on-premises, or fully air-gapped environment for sovereign deployments, an offering the company anchors on SOC 2 Type II and GDPR compliance. Layered on top is an Agent SDK for routing, tools, scoped retrieval, correction memory, and audit trails, plus proprietary infrastructure tooling such as the EasyR1 reinforcement-learning fine-tuning framework (GRPO/DAPO on veRL) and the open-source FlexBench MLPerf-style LLM inference benchmark.
The business model is API-first and product-led. Developers sign up at tokens.flex.ai, receive $10/month in free credits for three months, and are billed per token on serverless workloads or per-hour on dedicated GPU capacity. Distribution spans self-serve (tokens.flex.ai), enterprise sales for Dedicated Endpoints and AI Factory, and channel/reseller relationships through NVIDIA Inception, Microsoft for Startups, AWS Startups, Google for Startups, and a System Integrators program. Customer segments are split between individual builders/indie developers, growing AI teams needing predictable compute, and regulated enterprises requiring data residency, with selected reference customers including DragonLLM (sovereign finance inference), LegML (domain-specific legal LLM training), and Pixelcut (image generation fine-tuning).
FlexAI firmographics
Firmographics- Name
- FlexAI
- Legal name
- FlexAI SAS
- Website
- https://flex.ai
- Company type
- Private
- Founded year
- 2023
- Operating status
- Operating
- Headcount range
- 11–50 employees
- Short description
- FlexAI is a Paris-based AI infrastructure company offering an OpenAI-compatible platform for running 30+ open-weight models on heterogeneous NVIDIA/AMD GPUs across multiple clouds, with serverless, dedicated, and private-cloud deployment tiers for builders, growing AI teams, and regulated enterprises.
- Ownership category
- akta.pro rank
FlexAI industry classification
Industry- Product category
- AI Cloud Infrastructure
- NAICS
- Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (5182)
- SIC
- Services-Prepackaged Software (7372)
- akta.pro primary industry
- Model Serving, Inference & Deployment Platforms (APIs, Edge/On-Prem) (HDAEANAC)
- akta.pro secondary industries
- Model Hosting, Serving & Inference Platforms (HDAAACAB), Model Deployment, Serving & Inference Platforms (HDAAABAF), AI Application Enablement Platforms (Copilot/Agent Frameworks, SDKs) (HDAEANAJ)
Keywords
Where FlexAI is headquartered
LocationHeadquarters
- HQ city
- Paris
- HQ country
- France
- HQ region
- Europe
Offices3 records
Markets served
FlexAI business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Infrastructure, Technology or R&D, Personnel, Operations, Marketing or Sales
Revenue model
- Token Factory (Serverless Inference): Pay-per-token usage for serverless model inference via OpenAI-compatible API. Per-second billing with no idle charges. Market-tracked pricing linked to source providers.
- Dedicated Endpoints: Reserved GPU capacity (H100/H200) with flat hourly rates. Predictable monthly costs for production workloads with autoscaling and scale-to-zero capabilities.
- AI Factory (Private Cloud): Enterprise deployment in VPC, on-premises, or air-gapped environments with custom infrastructure arrangements.
- Training & Fine-tuning Services: GPU compute for distributed training and reinforcement learning workloads with preprovisioned environments and managed job orchestration.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Freemium | Monthly | Free Credits: $10/month for first 3 months |
| Usage-based | Pay-as-you-go | Serverless (Token Factory): Pay per token |
| Usage-based | Pay-as-you-go | Dedicated Endpoints: Reserved GPU capacity |
| Subscription | Annual | Annual Subscription: Committed capacity |
Go-to-market motion2 records
Distribution channels3 records
Marketing channels6 records
FlexAI product offering
Product offeringCore offering
FlexAI operates an agent-native AI compute platform that provides a single OpenAI-compatible API key for serverless inference across 30+ open models (text, vision, image, audio, video) running on heterogeneous NVIDIA and AMD GPUs. The platform extends from pay-per-token serverless usage to reserved Dedicated Endpoints and fully managed private AI cloud deployments (VPC, on-premises, or air-gapped) for sovereign workloads, and includes the Agent SDK for building production agent loops.
Product overview
FlexAI is an AI compute company offering a unified platform for agent-native AI infrastructure with three core product tiers: Token Factory (serverless model APIs), Agent SDK (agent development framework), and AI Factory (private cloud deployment). The platform enables AI model deployment across multiple hardware architectures (NVIDIA, AMD) with OpenAI-compatible APIs, supporting inference, fine-tuning, and training workloads. Additional modules include Dedicated Endpoints for reserved GPU capacity, pre-configured agent pipelines for coding, research, support, workflow automation, and multimodal generation, plus tools like EasyR1 for reinforcement learning fine-tuning and FlexBench for benchmarking. The portfolio ranges from pay-per-token serverless access to dedicated infrastructure with data sovereignty options.
Differentiator
Problem solved
Functional benefit
Brands
- Token Factory: OpenAI-compatible API providing access to 30+ models for text, vision, image, video, and audio.
- Agent SDK
- AI Factory
- Dedicated Endpoints
Products and services
- Token Factory Serverless AI inference product providing OpenAI-compatible API access to 30+ open models for text, vision, image, audio, and video generation under a single API key, with pay-per-token usage-based billing. Targeted at developers, startups, and enterprise teams that need flexible multi-model access without infrastructure commitment.
- Agent SDK SDK for building production agent loops with tool calling, streaming, structured output, and JSON/vision support, including routing, governance, scoped access, retrieval, tools, correction memory, and audit trails. Targeted at AI engineers building agentic applications on FlexAI infrastructure.
- AI Factory Managed private AI cloud deployment option delivered into VPC, on-premises, or air-gapped environments, providing sovereignty and compliance for regulated industries with autoscaling and scale-to-zero. Targeted at large enterprises in finance, healthcare, and government with sovereignty requirements.
- Dedicated Endpoints Reserved NVIDIA H100/H200 GPU instances providing dedicated compute capacity for production AI workloads, exposed through the same OpenAI-compatible API as Token Factory, with autoscaling, scale-to-zero, and per-second billing. Targeted at mid-market and enterprise teams with predictable production inference needs.
- Inference Managed inference service running customer-deployed models across heterogeneous NVIDIA and AMD GPU fleets with up to 99.9% uptime SLA. Targeted at enterprise teams requiring managed multi-vendor model serving.
- Fine-Tuning Managed fine-tuning service for customizing open-weight models on proprietary data, with usage-based pricing and managed GPU provisioning, job scheduling, and checkpoint management. Targeted at teams adapting open models to domain-specific tasks.
- Training Managed distributed training infrastructure providing preprovisioned environments, DDP/FSDP/ZeRO execution with bf16/fp8 mixed precision, and preemption-aware job management with sharded checkpointing and idempotent resume. Targeted at ML teams training or fine-tuning large models on GPU clusters.
- EasyR1 Open reinforcement learning fine-tuning framework supporting GRPO and DAPO algorithms for reasoning-focused post-training, built on the veRL training stack for distributed rollouts. Targeted at ML researchers and engineers applying RL fine-tuning to reasoning models on FlexAI infrastructure.
- FlexBench Open-source modular MLPerf-style LLM inference benchmark for measuring serving performance and evaluating open-source models across hardware configurations. Targeted at ML engineers and infra teams selecting hardware and quantifying serving performance.
Quantifiable outcome
- 70% cost reduction: 8x H100 on Azure ($20,102/mo) vs. FlexAI ($6,048/mo)
- +5 more outcomes
Companies that use FlexAI
Customer profileNamed customers3 records
Segments4 records
Ideal customer profiles4 records
FlexAI technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
AI capability8 records
Feature9 records
FlexAI partnerships and signals
Strategic signalPartnerships
Eleven partnerships are on record, tiered program_membership, hardware_partner, cloud_partner and core_partner.
- NVIDIA Inception Programprogram_membershipNVIDIA's accelerator program for AI startups providing technical support, tools, and ecosystem connections.
- Microsoft for Startupsprogram_membershipStartup accelerator program providing Azure credits, technical support, and go-to-market resources.
- AWS Startupsprogram_membershipAWS startup program providing credits, technical support, and access to AWS infrastructure.
- Google for Startupsprogram_membershipGoogle's startup accelerator program providing cloud credits and technical mentorship.
- NVIDIAhardware_partnerHardware provider for FlexAI's GPU fleet (H100, H200, B200). FlexAI works with NVIDIA as part of heterogeneous compute strategy.
- AMDhardware_partnerAlternative GPU provider for FlexAI's heterogeneous compute infrastructure alongside NVIDIA.
- AWScloud_partnerCloud provider partner for FlexAI's multi-cloud heterogeneous compute strategy.
- Google Cloudcloud_partnerCloud provider partner for FlexAI's multi-cloud heterogeneous compute strategy.
- Intelhardware_partnerHardware partner referenced in FlexAI's heterogeneous compute approach alongside NVIDIA and AMD.
- Compute Partnerscore_partnerGPU capacity contributors behind FlexAI's managed scheduling. FlexAI brokers demand across customer base. Relationships confidential by default.
- Model Providerscore_partnerAI model providers getting models served, priced, and distributed through Token Factory with readiness pipeline and per-model pricing.
Scale indicators5 records
Recent moves6 records
Expansion highlights7 records
FlexAI competitors and assessment
Company assessmentDirect peers
- Together AI: Together AI operates an OpenAI-compatible inference cloud serving 200+ open models across NVIDIA GPUs with similar per-token pricing and dedicated endpoints — the closest direct competitor to FlexAI's Token Factory proposition.
- Fireworks AI: Fireworks AI provides serverless and dedicated LLM inference with OpenAI-compatible APIs across open models and proprietary fine-tunes, competing head-to-head with FlexAI's Token Factory and Dedicated Endpoints tiers.
- Replicate: Replicate runs a cloud API for open-source ML models including LLMs, image, and audio models on a per-prediction basis — a directly comparable developer-facing inference marketplace to FlexAI.
- Anyscale: Anyscale (built on Ray) offers managed compute for AI workloads including LLM serving and distributed training, targeting the same developer and enterprise workloads as FlexAI's training and inference products.
- Modal Labs: Modal provides serverless GPU compute for AI/ML workloads with per-second billing and developer-first APIs — closely mirrors FlexAI's serverless-to-dedicated product ladder and target developer persona.
- Lambda Labs: Lambda operates GPU cloud infrastructure with on-demand and reserved H100 instances plus a hosted inference product, competing with FlexAI's Dedicated Endpoints and AI Factory offerings for training and inference workloads.
- CoreWeave: CoreWeave is a large-scale GPU cloud provider serving AI training and inference workloads on NVIDIA hardware — overlaps with FlexAI on dedicated GPU capacity and is a comparable reference point for the GPU-cloud business model.
Emerging players
- Mistral AI: Paris-based Mistral develops open-weight LLMs and a hosted inference API (La Plateforme), competing with FlexAI on European sovereign AI positioning and overlapping model coverage on Llama-class open models.
- Groq: Groq offers low-latency LLM inference on its custom LPU hardware via a cloud API, competing with FlexAI's serverless inference tier on price-performance for open-weight models.
Broad incumbents
- AWS Bedrock: AWS Bedrock is the hyperscaler incumbent offering serverless access to multiple foundation models with enterprise compliance and VPC deployment — overlaps with FlexAI's Token Factory and AI Factory for regulated enterprise buyers.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat5 records
Key risks6 records
Key highlights7 records
Customer concentration
FlexAI social profiles
Digital presenceFlexAI compliance and trust
Trust signalCompliance2 records
FlexAI financial estimates
Financial estimateRevenue estimate
Valuation estimate
FlexAI leadership team
Management profileNumber of profiles
Profiles2 records
FlexAI funding detail
Funding detailFunding overview
Funding rounds1 record
Investors8 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
FlexAI M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about FlexAI
What does FlexAI do?
FlexAI operates an agent-native AI compute platform that provides a single OpenAI-compatible API key for serverless inference across 30+ open models (text, vision, image, audio, video) running on heterogeneous NVIDIA and AMD GPUs. The platform extends from pay-per-token serverless usage to reserved Dedicated Endpoints and fully managed private AI cloud deployments (VPC, on-premises, or air-gapped) for sovereign workloads, and includes the Agent SDK for building production agent loops.
Is FlexAI a public or private company?
FlexAI is a private company. It is classified as venture growth investor backed and is currently operating.
When was FlexAI founded?
FlexAI was founded in 2023. It employs 11 to 50 people.
Where is FlexAI based?
FlexAI is headquartered in Paris, France, in the Europe region.
How does FlexAI make money?
Four revenue lines are on record. Token Factory (Serverless Inference) is the primary driver. The others are dedicated Endpoints, AI Factory (Private Cloud) and training & Fine-tuning Services.
Who are FlexAI's main competitors?
Direct peers on record are Together AI, Fireworks AI, Replicate, Anyscale, Modal Labs, Lambda Labs and CoreWeave. Emerging players are Mistral AI and Groq. AWS Bedrock is listed as a broad incumbent.
Does FlexAI have an API?
Yes. OpenAI-compatible HTTP API providing access to 30+ open models for text, vision, image generation, video, and audio. One API key for all models. Supports chat completions, streaming, tool calls, structured output, and vision inputs. Documentation available at docs.flex.ai with Python SDK and CLI reference. Developer documentation is at docs.flex.ai.
What industry is FlexAI in?
FlexAI's product category is AI Cloud Infrastructure. Its primary akta.pro industry code is HDAEANAC, Model Serving, Inference & Deployment Platforms (APIs, Edge/On-Prem), with a secondary code of HDAAACAB, Model Hosting, Serving & Inference Platforms. Its NAICS code is 5182 and its SIC code is 7372.