Modal Labs
Modal Labs provides a serverless GPU cloud platform purpose-built for AI workloads, offering inference, training, batch processing, and sandboxed code execution via a Python SDK. The platform serves AI/ML developers, AI-native startups (Suno, Runway, Decagon, Lovable), and enterprises (Quora) with usage-based pricing and self-service onboarding.
- Company typePrivate
- Founded2021
- HeadquartersNew York, United States
- Headcount51–100
- GTM typeB2B
- OfferingSoftware
What Modal Labs does
Modal Labs operates a serverless GPU cloud platform purpose-built for AI workloads. The platform provides an AI-native runtime with sub-second container cold starts, instant autoscaling from zero to over 1,000 GPUs, and multi-cloud GPU routing across regions in real time. The product portfolio rests on a Core Platform (the underlying container runtime) and includes three core products — Inference (LLM and multi-modal inference serving with vLLM and OpenAI-compatible APIs), Training (single-GPU fine-tuning up to 128x B200 multi-node training with 3,200 Gbps Infiniband, RL, and parallel hyperparameter sweeps), and Sandboxes (isolated, ephemeral GPU environments for coding agents, background agents, and RL rollouts at hundreds of thousands of concurrent runs) — plus Batch and Notebooks modules. The entire platform is programmable from Python via the Modal SDK, with no YAML or infrastructure configuration files required.
The business model is usage-based, charging per second of GPU compute, with a $30 free credit for new sign-ups and a fully self-service product-led go-to-market. Customers span AI-native startups (Suno, Runway, Decagon, Physical Intelligence, Lovable, Chai Discovery, Reducto, Substack) and larger enterprises (Quora), with documented outcomes including 65% latency reduction (Decagon), 3x latency decrease (Reducto), 10-15 ms latency for robot control (Physical Intelligence), and 4 months faster time-to-launch (Suno). The platform is SOC2 Type 2 and HIPAA compliant, has reported 99.937% uptime over a 30-day window, and supports sub-10ms overhead latency for online inference from globally distributed compute. Modal's customer base is horizontal, and marketing and developer relations are built around documentation, an open-source GitHub examples repository, a Slack community, technical blog posts, a GPU Glossary, and YouTube tutorials.
Modal Labs firmographics
Firmographics- Name
- Modal Labs
- Legal name
- Modal Labs
- Website
- https://modal.com
- Company type
- Private
- Founded year
- 2021
- Operating status
- Operating
- Headcount range
- 51–100 employees
- Short description
- Modal Labs provides a serverless GPU cloud platform purpose-built for AI workloads, offering inference, training, batch processing, and sandboxed code execution via a Python SDK. The platform serves AI/ML developers, AI-native startups (Suno, Runway, Decagon, Lovable), and enterprises (Quora) with usage-based pricing and self-service onboarding.
- Ownership category
- akta.pro rank
Modal Labs industry classification
Industry- Product category
- Cloud GPU Infrastructure
- NAICS
- Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (518), Computer Facilities Management Services (541513), Computer Systems Design and Related Services (54151)
- SIC
- Services-Prepackaged Software (7372), Services-Computer Processing & Data Preparation (7374)
- akta.pro primary industry
- AI Compute Cloud & GPU-as-a-Service (HDAAAAAK)
- akta.pro secondary industries
- Serverless & Functions-as-a-Service (FaaS) (HDABAAAC), AI Compute Virtualization & Scheduling (GPU virtualization, cluster schedulers) (HDAAAAAG), GPU-Accelerated & AI Training/Inference Servers (HDACABAG)
Keywords
Where Modal Labs is headquartered
LocationHeadquarters
- HQ city
- New York
- HQ country
- United States
- HQ region
- North America
Markets served
Modal Labs business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Technology or R&D, Infrastructure, Personnel, Marketing or Sales, Operations
Revenue model
- GPU compute consumption: Pay-per-second usage for GPU compute time. Users scale from 0 to 1000+ GPUs and pay only for what they use. No reserved capacity required.
- Free tier / trial credit: $30 free compute credit for new users to try the platform. Enables product-led onboarding without upfront commitment.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Freemium | Monthly | Free tier with $30 compute credit |
Go-to-market motion1 record
Distribution channels1 record
Marketing channels6 records
Modal Labs product offering
Product offeringCore offering
Modal Labs operates an AI-native serverless GPU cloud platform that lets developers run inference, training, batch processing, and isolated code execution on NVIDIA GPUs (H100s, A100s, B200s, H200s, L40S, A10Gs) without managing infrastructure. Workloads are defined in Python via the Modal SDK, deploy in seconds, and elasticly scale from 0 to 1,000+ GPUs across multiple cloud providers with sub-second cold starts.
Product overview
Modal Labs offers an AI-native cloud platform for serverless GPU workloads. The Core Platform provides the foundational runtime with sub-second cold starts, instant autoscaling (0 to 1000+ GPUs), and GPU infrastructure (H100s, A100s, A10Gs, B200s, H200s, L40S). Built on top are three core products: Inference (LLM and multi-modal inference serving), Training (fine-tuning, multi-node training, RL, hyperparameter sweeps), and Sandboxes (isolated code execution for agents and untrusted code). Add-on modules include Batch (parallel batch inference), Notebooks (Jupyter on GPUs), and the Modal SDK (Python developer interface). The entire portfolio is programmable from Python with no YAML, no infrastructure configuration files, and scales to zero between requests. Pricing offers $30/month free compute, with GPU workloads billed per second.
Differentiator
Problem solved
Functional benefit
Brands
- Modal SDK: Python SDK for cloud compute infrastructure
- Modal Cloud
- DoppelBot
- GPU Glossary
Products and services
- Inference Serverless inference product for deploying and scaling LLM and multi-modal inference workloads (text, image, video, audio, embeddings). Supports any model or inference engine on H100s, A100s, A10Gs and more, with scale-to-zero between requests and burst scaling. Includes OpenAI-compatible API endpoints, token streaming, WebRTC, and WebSocket support.
- Training End-to-end training product for fine-tuning and training ML models, from single-GPU SFT, LoRA, and full fine-tunes to multi-node clusters up to 128 B200s with 3200 Gbps Infiniband networking. Supports reinforcement learning with thousands of concurrent trajectories and parallel hyperparameter sweeps.
- Sandboxes Isolated, ephemeral, GPU-accelerated execution environments for running untrusted code, coding agents, background agents, and RL rollouts. Spins up hundreds of thousands of concurrent rollout environments in seconds with custom images and dependencies.
- Batch
Quantifiable outcome
- 65% latency reduction for Decagon
- +5 more outcomes
Companies that use Modal Labs
Customer profileNamed customers9 records
Segments4 records
Ideal customer profiles2 records
Modal Labs technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration10 records
AI capability11 records
Feature9 records
Modal Labs partnerships and signals
Strategic signalPartnerships
Four partnerships are on record, tiered core.
- Hugging FacecoreModal integrates with Hugging Face Hub for model hosting and downloads. Modal Volumes are used for caching Hugging Face model weights. Users can fine-tune models from Hugging Face and deploy inference endpoints.
- vLLMcoreModal uses vLLM as the primary inference engine for LLM serving, with native support for LoRA adapter swapping and OpenAI-compatible APIs.
- AWS S3coreModal CloudBucketMount feature enables mounting S3 buckets directly into Modal apps for reading and writing cloud storage as local directories.
- NVIDIAcoreModal provides access to NVIDIA GPUs (H100s, A100s, A10Gs, B200s, H200s, L40S) across cloud providers. GPU scheduling and fleet management built on NVIDIA infrastructure.
Scale indicators5 records
Recent moves6 records
Expansion highlights7 records
Modal Labs competitors and assessment
Company assessmentDirect peers
- CoreWeave: CoreWeave is a GPU-focused cloud provider offering large-scale H100/B200 compute for AI training and inference. It is the closest direct peer to Modal — same customer base (AI labs, startups, enterprises), same product (GPU cloud), competing on performance, scale, and developer tooling.
- Lambda Labs: Lambda operates GPU cloud instances and clusters for AI training and inference, with both on-demand and reserved offerings. Competes with Modal for AI developers seeking high-performance GPU compute, with a similar focus on Python/ML workflows.
- Together AI: Together AI provides serverless and dedicated GPU inference for open-source LLMs, with a developer-friendly API and per-token pricing. Closely parallels Modal's Inference product in targeting AI startups with pay-as-you-go GPU inference.
- RunPod: RunPod offers on-demand GPU pods and serverless endpoints for AI inference and training at competitive prices. Targets the same developer-led AI workload segment as Modal and competes on price and per-second billing.
- Anyscale: Anyscale (creators of Ray) provides managed compute for distributed AI workloads, including training, fine-tuning, and inference at scale. Competes with Modal's Training product for ML teams running multi-node training and RL rollouts.
- Replicate: Replicate offers a serverless cloud for running open-source ML models via API, with a developer-focused, PLG motion similar to Modal. Targets AI developers wanting simple per-prediction inference without managing infrastructure.
Emerging players
- Fireworks AI: Fireworks AI provides optimized open-source LLM inference and fine-tuning APIs with low-latency serving. Comparable to Modal's Inference and LoRA-swap capabilities, with stronger emphasis on model API rather than raw GPU infrastructure.
Broad incumbents
- AWS (SageMaker / Bedrock): AWS offers managed AI/ML infrastructure (SageMaker) and foundation-model services (Bedrock) with broader enterprise reach. Competes with Modal across the same workloads but lacks Modal's AI-native developer ergonomics and sub-second cold-start focus.
- Google Cloud Vertex AI: Google Cloud's Vertex AI provides managed training, tuning, and inference on TPUs/GPUs with deep integration into the Google ecosystem. Competes broadly with Modal for AI workloads, particularly for customers already on GCP.
- Microsoft Azure AI: Azure AI offers managed compute, model serving, and MLOps tooling integrated with the broader Azure enterprise stack. Competes with Modal for enterprise AI workloads, particularly via OpenAI partnership and Azure ML.
Market position
Strengths5 records
Weaknesses4 records
Competitive moat5 records
Key risks5 records
Key highlights7 records
Customer concentration
Modal Labs social profiles
Digital presenceModal Labs compliance and trust
Trust signalCompliance2 records
Modal Labs financial estimates
Financial estimateRevenue estimate
Valuation estimate
Modal Labs leadership team
Management profileNumber of profiles
Profiles1 record
Modal Labs funding detail
Funding detailFunding overview
Funding rounds5 records
Investors19 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Modal Labs M&A and investment
M&A and investmentM&A2 records
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Modal Labs
What does Modal Labs do?
Modal Labs operates an AI-native serverless GPU cloud platform that lets developers run inference, training, batch processing, and isolated code execution on NVIDIA GPUs (H100s, A100s, B200s, H200s, L40S, A10Gs) without managing infrastructure. Workloads are defined in Python via the Modal SDK, deploy in seconds, and elasticly scale from 0 to 1,000+ GPUs across multiple cloud providers with sub-second cold starts.
Is Modal Labs a public or private company?
Modal Labs is a private company. It is classified as venture growth investor backed and is currently operating.
When was Modal Labs founded?
Modal Labs was founded in 2021. It employs 51 to 100 people.
Where is Modal Labs based?
Modal Labs is headquartered in New York, United States, in the North America region.
How does Modal Labs make money?
Two revenue lines are on record. GPU compute consumption is the primary driver. The others are free tier / trial credit.
Who are Modal Labs's main competitors?
Direct peers on record are CoreWeave, Lambda Labs, Together AI, RunPod, Anyscale and Replicate. Fireworks AI is listed as an emerging player. Broad incumbents are AWS (SageMaker / Bedrock), Google Cloud Vertex AI and Microsoft Azure AI.
Does Modal Labs have an API?
Yes. Modal provides a Python SDK (modal) that allows developers to define and deploy serverless functions and classes via decorators such as @app.function(), @app.cls(), @app.server(), @modal.wsgi_app(), @modal.asgi_app(), and @modal.fastapi_endpoint(). The SDK exposes methods for remote execution (.remote()), local execution (.local()), parallel mapping (.map(), .starmap()), spawning without waiting (.spawn()), batching (.batched), and dynamic configuration overrides (.with_options(), .with_concurrency(), .with_batching()). Functions can be invoked via the modal run CLI, modal deploy, or programmatically. Secrets management, volume mounts, GPU allocation, scheduling, and autoscaling are all configurable through the SDK. An OpenAI-compatible API is supported via vLLM serving for LLM inference. Developer documentation is at modal.com/docs.
What industry is Modal Labs in?
Modal Labs's product category is Cloud GPU Infrastructure. Its primary akta.pro industry code is HDAAAAAK, AI Compute Cloud & GPU-as-a-Service, with a secondary code of HDABAAAC, Serverless & Functions-as-a-Service (FaaS). Its NAICS code is 518 and its SIC code is 7372.