Modular
Modular is an AI infrastructure company building a unified inference platform (MAX framework) and high-performance programming language (Mojo) enabling AI models to run efficiently across NVIDIA, AMD, Apple Silicon, Intel, and ARM hardware. Acquired by Qualcomm for $3.92B in June 2026.
- Company typePrivate
- Founded2022
- HeadquartersPalo Alto, United States
- Headcount101–250
- GTM typeB2B
- OfferingSoftware
What Modular does
Modular is an AI infrastructure company founded in 2022 by Chris Lattner (creator of LLVM, Clang, Swift, and MLIR) and Tim Davis (Google AI infrastructure veteran behind TensorFlow, XLA, MLIR, and TFLite). The company develops a unified AI inference platform consisting of two core products: the MAX framework, a GenAI-native modeling and serving platform that automatically optimizes kernels and request execution across accelerators; and Mojo, a high-performance systems programming language built on MLIR compiler infrastructure that combines Python-like syntax with C-level performance for writing GPU kernels. The platform targets the CUDA lock-in problem by enabling hardware-portable AI deployment across 15+ architectures including NVIDIA, AMD, Apple Silicon, Intel, and ARM, claiming 2x performance improvement over vLLM on diverse hardware through a single container and OpenAI-compatible API.
Modular's business model combines usage-based pricing (per-token for shared endpoints, per-minute for dedicated NVIDIA and AMD GPUs) with subscription-based enterprise deployments (self-hosted containers and customer VPC options), supplemented by a freemium tier for developer acquisition. Go-to-market is hybrid: product-led growth through self-serve console signup, Discord/GitHub community engagement, and GPU Kernel Hackathons, combined with direct enterprise sales for organizations requiring SLAs and dedicated infrastructure. Customers span hyperscalers (AWS, Oracle Cloud Infrastructure), AI application companies (TensorWave, Inworld, Hippocratic AI, Luma AI), and a long tail of developers reached through PLG channels. Distribution includes the Modular Console (self-serve), direct enterprise sales, AWS Marketplace, and Oracle Cloud Infrastructure.
Modular raised $380 million across three funding rounds (June 2022 seed, August 2023 Series A, September 2025 growth round) reaching a $1.6 billion valuation, with investors including General Catalyst, GV, Greylock, DFJ Growth, and Thomas Tull's US Innovative Technology Fund. In June 2026, Qualcomm announced the acquisition of Modular for approximately $3.92 billion in an all-stock transaction, expected to close H2 2026. Prior to the acquisition, the company expanded globally with offices in Edinburgh and San Francisco (April 2026), acquired BentoML to broaden deployment capabilities (February 2026), and formed strategic partnerships with AMD, Oracle, and Hugging Face to deepen ecosystem reach across heterogeneous compute.
Modular firmographics
Firmographics- Name
- Modular
- Legal name
- Modular Inc.
- Website
- https://modular.com
- Company type
- Private
- Founded year
- 2022
- Operating status
- Acquired
- Headcount range
- 101–250 employees
- Short description
- Modular is an AI infrastructure company building a unified inference platform (MAX framework) and high-performance programming language (Mojo) enabling AI models to run efficiently across NVIDIA, AMD, Apple Silicon, Intel, and ARM hardware. Acquired by Qualcomm for $3.92B in June 2026.
- Ownership category
- akta.pro rank
Modular industry classification
Industry- Product category
- AI Inference Infrastructure
- NAICS
- Software Publishers (5132), Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (5182)
- SIC
- Services-Prepackaged Software (7372), Services-Computer Programming Services (7371)
- akta.pro primary industry
- Model Deployment, Serving & Inference Platforms (HDAAABAF)
- akta.pro secondary industries
- AI Compiler, Runtime & Kernel Optimization Software (CUDA/ROCm/XLA, graph compilers) (HDAAAAAI), AI Server Systems & HGX/Accelerator Platforms (HDAAAAAB), AI Integration & Orchestration Platforms (Connectors, Workflow, iPaaS for AI) (HDAEANAI)
Keywords
Where Modular is headquartered
LocationHeadquarters
- HQ city
- Palo Alto
- HQ country
- United States
- HQ region
- North America
Offices3 records
Markets served
Modular business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Personnel, Technology or R&D, Infrastructure, Marketing or Sales, Operations
Revenue model
- Managed Cloud Inference: Fully managed inference service running on Modular's hosted cloud infrastructure with NVIDIA and AMD GPUs. Customers pay for usage based on token consumption or minute-based billing. Includes SLAs and security managed by Modular.
- Dedicated Endpoints: Reserved GPU instances (NVIDIA and AMD) with dedicated resources for mission-critical reliability. Per-minute pricing with flexible billing.
- Custom Model Deployment: Bring-your-own custom or fine-tuned models deployed on optimized infrastructure with per-minute pricing.
- Enterprise Software Licenses: Enterprise deployments in customer VPCs ('Your Cloud' option) with Modular managing the control plane while inference runs in the customer's infrastructure.
- Self-Hosted MAX: MAX and Mojo available as containerized software for self-hosting on customer infrastructure (NVIDIA, AMD, and other supported hardware).
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Freemium | Monthly | Free tier for development and evaluation |
| Usage-based | Pay-as-you-go | Shared Endpoints - Pay per token |
| Usage-based | Pay-as-you-go | Dedicated Endpoints - Reserved GPUs |
| Usage-based | Pay-as-you-go | Custom Models - Bring your own |
| Subscription | Annual | Enterprise - Contact sales |
Go-to-market motion2 records
Distribution channels6 records
Marketing channels8 records
Modular product offering
Product offeringCore offering
Modular develops and sells a unified AI inference platform built around the MAX framework and the Mojo programming language. The platform enables deployment, serving, and programming of AI models across heterogeneous compute (NVIDIA, AMD, Intel, ARM, and Apple Silicon GPUs and CPUs) with full-stack optimizations from GPU kernel to API endpoint, delivered via fully managed cloud, customer VPC deployment, and self-hosted containers.
Product overview
Modular offers a unified AI inference platform combining the MAX framework and Mojo programming language. MAX provides genAI-native modeling and serving with automatic kernel optimization across diverse hardware (NVIDIA, AMD, Intel, ARM, Apple Silicon), delivering 2x performance over vLLM. Mojo is a high-performance systems language for GPU/CPU kernels achieving Python-like simplicity with near-C speed. The platform supports text, image, video, audio, and code generation through shared, dedicated, or custom model endpoints. Deployment options include Modular's fully managed cloud, customer VPC deployment, and self-hosted containers. BentoML joined Modular in February 2026, expanding deployment capabilities.
Differentiator
Problem solved
Functional benefit
Brands
- MAX: A unified AI inference platform for high-performance, portable compute enabling full optimizations from GPU kernel to API endpoint.
- Mojo
Products and services
- MAX Framework A unified GenAI-native modeling and serving framework providing hardware-agnostic, high-performance inference with automatic kernel optimization. Delivers 2x performance improvement over vLLM through a single container with OpenAI-compatible API. Supports 1000+ models out of the box.
- Mojo Programming Language A high-performance systems programming language for AI that combines Python-like syntax with systems-level control, enabling developers to write custom GPU kernels for CPUs and GPUs with performance approaching C/C++. Built on MLIR compiler infrastructure.
- Modular Cloud Fully managed cloud inference service operated by Modular on NVIDIA and AMD GPUs, providing access to frontier models via API with per-token or per-minute pricing and no infrastructure management required.
- Your Cloud (VPC Deployment) Deployment of the Modular stack in a customer's Virtual Private Cloud, where Modular manages the control plane while inference runs in the customer's own infrastructure. Customer owns hardware, data, and cloud credits.
- Self-Hosted MAX Containerized software deployment of MAX and Mojo on customer infrastructure, supporting NVIDIA, AMD, and other hardware with customer-controlled operational policies.
- Shared Endpoints High-performance inference service with no infrastructure management, providing per-token pricing and access to frontier models including DeepSeek V3, Kimi K2.6, and MiniMax M3.
- Dedicated Endpoints Reserved NVIDIA and AMD GPUs for mission-critical reliability, offering per-minute pricing with predictable performance and dedicated resource allocation for production workloads.
- Custom Models Deployment service for bringing custom or fine-tuned models onto Modular's optimized infrastructure with per-minute pricing and forward-deployed engineering support for porting.
- Mojo Agent Skills Official AI agent skills from Modular available on GitHub that enable AI agents to perform automated tasks using MAX and Mojo capabilities across different hardware platforms.
Quantifiable outcome
- 2x performance improvement over vLLM on diverse hardware
- +4 more outcomes
Companies that use Modular
Customer profileNamed customers11 records
Segments4 records
Ideal customer profiles3 records
Modular technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration2 records
AI capability12 records
Feature6 records
Modular partnerships and signals
Strategic signalPartnerships
Five partnerships are on record, tiered strategic.
- Hugging FacestrategicExpanded strategic partnership announced June 2026 at Qualcomm Investor Day. Partnership focuses on three pillars: connecting Qualcomm's data center infrastructure with Hugging Face's AI storage, accelerating AI model deployment across devices, and enabling agentic AI orchestration across hybrid environments. Targets Hugging Face's 16 million developers.
- SqueezeBitsstrategicSouth Korean AI optimization startup partnered with Modular to expand MAX inference framework into generative media. SqueezeBits leads expansion into Korean and Asian markets with combined solution of model compression and high-speed performance. Partnership was previewed at NVIDIA GTC 2026.
- Oracle Cloud Infrastructure (OCI)strategicModular selected Oracle Cloud Infrastructure as preferred platform for AI startups and enterprises. Oracle provides high-performance GPU clusters, low-latency networking, and reliable global infrastructure for Modular's customers.
- AMDstrategicDeep partnership for AMD GPU optimization. Modular + AMD initiative 'Unleashing AI Performance on AMD GPUs' enables state-of-the-art inference performance on AMD MI355 and other AMD accelerators. Joint participation at AMD AI DevDay 2026.
- InworldstrategicCustomer case study partnership demonstrating AI character voice generation powered by Modular MAX. Partnership showcased at Oracle events.
Scale indicators15 records
Recent moves7 records
Expansion highlights6 records
Modular competitors and assessment
Company assessmentDirect peers
- OctoAI (acquired by NVIDIA): OctoAI provided fine-tuning and serving infrastructure for open models with hardware-portability emphases, positioning directly against MAX before its acquisition by NVIDIA.
- DeepInfra: DeepInfra runs serverless inference for open models on heterogeneous hardware with low per-token pricing, competing head-on with Modular's Shared Endpoints.
- Replicate: Replicate hosts open-source and custom AI models behind a cloud API with usage-based billing, directly comparable to Modular's pay-per-token inference model across modalities.
- Together AI: Together AI runs an OpenAI-compatible inference cloud for open-weight and custom models on heterogeneous GPUs, directly comparable to Modular's MAX-powered 'Shared Endpoints' and dedicated GPU offerings.
- Fireworks AI: Fireworks AI provides serverless and dedicated LLM inference APIs with custom optimizations, competing head-on with Modular on per-token pricing, OpenAI-compatible APIs, and fine-tuned model deployment.
- Anyscale: Anyscale (commercializing Ray) offers a distributed compute and serving platform for AI workloads, overlapping with Modular's model serving and scaling capabilities across GPUs.
Broad incumbents
- CoreWeave: CoreWeave is a large-scale GPU cloud provider that also layers inference services, overlapping with Modular on cloud-based and dedicated GPU deployment of AI workloads.
- Hugging Face: Hugging Face hosts thousands of models and now offers Inference API and dedicated endpoints, overlapping directly with Modular's inference catalog and a recently expanded strategic partnership.
- Lambda Labs: Lambda provides GPU cloud and inference APIs (Lambda Chat, Inference API), competing with Modular's managed cloud and dedicated endpoint products.
Emerging players
- RunPod: RunPod offers GPU cloud and serverless inference for AI models, with comparable usage-based pricing and BYO model deployment to Modular's Self-Hosted and Shared Endpoints.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat5 records
Key risks6 records
Key highlights7 records
Customer concentration
Modular social profiles
Digital presenceModular compliance and trust
Trust signalCompliance3 records
Modular financial estimates
Financial estimateRevenue estimate
Valuation estimate
Modular leadership team
Management profileNumber of profiles
Profiles6 records
Modular subsidiaries and ownership
Company hierarchySubsidiaries1 record
Modular funding detail
Funding detailFunding overview
Funding rounds3 records
Investors8 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Modular M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Modular
What does Modular do?
Modular develops and sells a unified AI inference platform built around the MAX framework and the Mojo programming language. The platform enables deployment, serving, and programming of AI models across heterogeneous compute (NVIDIA, AMD, Intel, ARM, and Apple Silicon GPUs and CPUs) with full-stack optimizations from GPU kernel to API endpoint, delivered via fully managed cloud, customer VPC deployment, and self-hosted containers.
Is Modular a public or private company?
Modular is a private company. It is classified as corporate owned and is currently acquired.
When was Modular founded?
Modular was founded in 2022. It employs 101 to 250 people.
Where is Modular based?
Modular is headquartered in Palo Alto, United States, in the North America region.
How does Modular make money?
Five revenue lines are on record. Managed Cloud Inference is the primary driver. The others are dedicated Endpoints, custom Model Deployment, enterprise Software Licenses and self-Hosted MAX.
Who are Modular's main competitors?
Direct peers on record are OctoAI (acquired by NVIDIA), DeepInfra, Replicate, Together AI, Fireworks AI and Anyscale. Broad incumbents are CoreWeave, Hugging Face and Lambda Labs. RunPod is listed as an emerging player.
Does Modular have an API?
Yes. Modular provides an OpenAI-compatible inference API that allows developers to access various AI models (including DeepSeek, Kimi K2.5, MiniMax M3, and custom models) through standard API calls. The API uses a base URL format and supports per-token and per-minute pricing models. The platform offers shared endpoints for non-dedicated access, dedicated endpoints with reserved NVIDIA and AMD GPUs, and custom model deployment options. Developer documentation is at docs.modular.com.
What industry is Modular in?
Modular's product category is AI Inference Infrastructure. Its primary akta.pro industry code is HDAAABAF, Model Deployment, Serving & Inference Platforms, with a secondary code of HDAAAAAI, AI Compiler, Runtime & Kernel Optimization Software (CUDA/ROCm/XLA, graph compilers). Its NAICS code is 5132 and its SIC code is 7372.