Luminal
Luminal is a San Francisco-based AI infrastructure company building a compiler-first inference platform that converts PyTorch and HuggingFace models into optimized native code for GPUs and ASICs, sold via self-serve usage-based cloud and licensed enterprise on-prem deployments.
- Company typePrivate
- Founded2025
- HeadquartersSan Francisco, United States
- Headcount11–50
- GTM typeB2B
- OfferingSoftware
What Luminal does
Luminal (Luminal AI Inc.) is a privately held, San Francisco-based AI infrastructure company founded in 2025 that builds a compiler-first AI inference platform targeting GPU and ASIC accelerators. The company's core product is the Luminal Compiler, which converts PyTorch and HuggingFace models ahead of time into optimized native code through a graph-level intermediate representation, hardware-aware optimization passes (fusion, tiling, memory planning, and scheduling), and zero-overhead code generation that emits directly to GPU kernels or ASIC instructions. This compiler-first approach contrasts with runtime inference engines (e.g., vLLM, TensorRT-LLM, native PyTorch), which interpret models dynamically and incur runtime overhead. The company pairs the compiler with the Luminal Inference OS, a hyperscale orchestration layer that schedules and load-balances inference workloads across heterogeneous compute (CPUs, GPUs, ASICs) with real-time dynamic load balancing and elastic node scaling.
Luminal operates two distribution channels: Luminal Cloud, a self-serve usage-based serverless inference service with scale-to-zero, automatic batching, and optimized compilation; and On-Prem Deployment, a licensed subscription product sold through direct enterprise engagement with dedicated engineering support, custom kernel optimization, and tailored SLAs. The company's interactive calculator illustrates an example monthly cost of $14.6k for an 8xH100 deployment serving 100,000 requests/day and 4.6B tokens, compared to $19.2k for OpenAI GPT-5 and $32.3k for Anthropic Claude Sonnet at the same volume. Self-reported benchmarks claim 3.2x throughput over vLLM, 1.3x over TensorRT-LLM, and 12x over native PyTorch on GPT-OSS 120B at 8xH100, with sub-10ms p99 latency.
The company is led by co-founder and CEO Joe Fioti, co-founder and CTO Matthew Gunton, and COO Jake Stevens, with a lean 11-50 person team working in person in San Francisco. Luminal raised $5.3 million in a seed round led by Felicis Ventures in November 2025, with participation from Liquid 2 Ventures, Saga Ventures, Palm Drive Capital, Y Combinator, and angel investors Paul Graham, Guillermo Rauch, and Ben Porterfield; an earlier $500K investment from Y Combinator was reported in September 2025. No customers, revenue, patents, or third-party benchmark validations are disclosed.
Luminal firmographics
Firmographics- Name
- Luminal
- Legal name
- Luminal AI Inc.
- Website
- https://luminal.com
- Company type
- Private
- Founded year
- 2025
- Operating status
- Operating
- Headcount range
- 11–50 employees
- Short description
- Luminal is a San Francisco-based AI infrastructure company building a compiler-first inference platform that converts PyTorch and HuggingFace models into optimized native code for GPUs and ASICs, sold via self-serve usage-based cloud and licensed enterprise on-prem deployments.
- Ownership category
- akta.pro rank
Luminal industry classification
Industry- Product category
- AI Inference Platform / ML Compiler Infrastructure
- NAICS
- Software Publishers (513210), Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (5182), Custom Computer Programming Services (541511)
- SIC
- Services-Prepackaged Software (7372), Services-Computer Programming, Data Processing, Etc. (7370)
- akta.pro primary industry
- Model Serving, Inference & Deployment Platforms (APIs, Edge/On-Prem) (HDAEANAC)
- akta.pro secondary industries
- Model Deployment, Serving & Inference Platforms (HDAAABAF), Model Hosting, Serving & Inference Platforms (HDAAACAB)
Keywords
Where Luminal is headquartered
LocationHeadquarters
- HQ city
- San Francisco
- HQ country
- United States
- HQ region
- North America
Offices1 record
Markets served
Luminal business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Technology or R&D, Personnel, Infrastructure, Operations, Marketing or Sales
Revenue model
- Luminal Cloud - Serverless Inference: Managed serverless inference service with automatic scaling. Users deploy in minutes and pay only for what they use. Features include serverless inference endpoints, scale-to-zero capabilities, automatic batching, and optimized compilation.
- On-Prem Deployment: Licensed deployment option for customers running Luminal on their own infrastructure. Includes dedicated engineering support, custom kernel optimization, and strict SLAs tailored to the customer. Can be deployed as cloud or on-prem.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Usage-based | Pay-as-you-go | Luminal Cloud - Pay-as-you-go serverless inference |
| Subscription | Multi-year contract | On-Prem Deployment - Licensed infrastructure |
Go-to-market motion1 record
Distribution channels2 records
Marketing channels5 records
Luminal product offering
Product offeringCore offering
Luminal provides a compiler-first AI inference platform that compiles PyTorch and Hugging Face models ahead of time into optimized native code for GPUs and ASICs, eliminating runtime interpretation overhead. Its Hyperscale Inference OS dynamically schedules and load-balances workloads across heterogeneous compute, and the platform is delivered via a managed serverless cloud offering (Luminal Cloud) or licensed on-prem deployment with dedicated engineering support.
Product overview
Luminal is a compiler-first AI inference platform consisting of two core products: the Luminal Compiler (which compiles PyTorch/Hugging Face models into optimized native code for GPUs and ASICs via graph-level IR, hardware-aware optimization, and zero-overhead codegen) and the Luminal Inference OS (which orchestrates heterogeneous compute workloads at scale with dynamic scheduling and load-balancing). The company offers two deployment options: Luminal Cloud (managed serverless inference with pay-per-use pricing) and On-Prem Deployment (licensed deployment with dedicated support). The platform consistently outperforms existing inference engines by 2-3x on standard benchmarks.
Differentiator
Problem solved
Functional benefit
Products and services
- Luminal Compiler Compiler that transforms PyTorch and Hugging Face models into optimized native code for GPUs and ASICs through graph-level IR lowering, hardware-aware optimization passes (fusion, tiling, memory planning, scheduling), and zero-overhead code generation.
- Luminal Inference OS Hyperscale inference orchestration layer that dynamically schedules and load-balances inference workloads across heterogeneous compute (CPUs, GPUs, ASICs) from single accelerators to large clusters, minimizing latency and maximizing throughput.
- Luminal Cloud Managed serverless inference platform that enables deployment in minutes with automatic scaling, scale-to-zero capabilities, automatic batching, and optimized compilation. Customers pay only for what they use.
- On-Prem Deployment Licensed deployment option that runs Luminal on customer infrastructure, with dedicated engineering support, custom kernel optimization, strict SLAs, and enterprise-grade security.
Quantifiable outcome
- 3.2x throughput improvement vs vLLM on GPT-OSS 120B, 8xH100 benchmark (Luminal 36k tok/s vs vLLM 26k tok/s)
- +4 more outcomes
Companies that use Luminal
Customer profileSegments1 record
Ideal customer profiles2 records
Luminal technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- No
- API docs
- API detail
Core technology
AI maturity
App detail
AI capability2 records
Feature7 records
Luminal partnerships and signals
Strategic signalScale indicators1 record
Recent moves5 records
Expansion highlights6 records
Luminal competitors and assessment
Company assessmentDirect peers
- vLLM (project): Open-source high-throughput LLM inference engine (UC Berkeley origin) that Luminal directly benchmarks against (3.2x advantage claim). Both target the same ML engineering audience deploying open-source LLMs at scale and compete on tokens/sec and latency.
- Fireworks AI: Cloud inference platform purpose-built for fast open-source LLM serving with custom kernels and a serverless deployment model. Closely overlaps Luminal's PLG cloud + on-prem dual GTM and competes for the same ML infrastructure budget.
- Together AI: Cloud provider for open-source LLM inference and training with proprietary optimization stack (e.g., Together Inference Engine). Competes head-to-head with Luminal Cloud on tokens/sec-per-dollar for open-source models.
- Anyscale: Ray-based platform for production AI workloads including large-scale LLM inference. Overlaps Luminal's hyperscale inference OS positioning (dynamic scheduling, load balancing across heterogeneous compute) for enterprise AI deployments.
- Modal: Serverless GPU compute platform with auto-scaling inference and developer-first deployment. Closely matches Luminal Cloud's PLG, serverless, scale-to-zero model for deploying AI workloads on demand.
- Replicate: Cloud API for running open-source ML models with usage-based pricing and automatic scaling. Competes for the same self-serve developer audience deploying Hugging Face / PyTorch models without managing infrastructure.
- Hugging Face Text Generation Inference (TGI): Open-source Rust-based inference server for LLMs maintained by Hugging Face, embedded in the dominant model hub. Competes with Luminal for the same PyTorch/HuggingFace model deployment workflow at enterprise and self-serve tiers.
Broad incumbents
- TensorRT-LLM (NVIDIA): NVIDIA's official LLM inference optimization library, with first-party access to GPU internals and tight CUDA integration. Directly benchmarked by Luminal (1.3x advantage) and the most strategically dangerous competitor due to NVIDIA's hardware/software co-design.
- OctoAI (acquired by NVIDIA): Cloud inference optimization platform (now part of NVIDIA) that historically targeted the same cost-efficient LLM serving use case. Demonstrates the strategic value of inference-optimization startups to GPU incumbents.
Emerging players
- DeepSpeed (Microsoft): Microsoft's open-source deep learning optimization library with inference components (DeepSpeed-MII, FasterTransformer) targeting large-model serving. Competes on the same axis of compiler/runtime optimization for transformer inference.
Market position
Strengths4 records
Weaknesses4 records
Competitive moat3 records
Key risks6 records
Key highlights6 records
Customer concentration
Luminal social profiles
Digital presenceLuminal financial estimates
Financial estimateRevenue estimate
Valuation estimate
Luminal leadership team
Management profileNumber of profiles
Profiles3 records
Luminal funding detail
Funding detailFunding overview
Funding rounds2 records
Investors6 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Luminal M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Luminal
What does Luminal do?
Luminal provides a compiler-first AI inference platform that compiles PyTorch and Hugging Face models ahead of time into optimized native code for GPUs and ASICs, eliminating runtime interpretation overhead. Its Hyperscale Inference OS dynamically schedules and load-balances workloads across heterogeneous compute, and the platform is delivered via a managed serverless cloud offering (Luminal Cloud) or licensed on-prem deployment with dedicated engineering support.
Is Luminal a public or private company?
Luminal is a private company. It is classified as venture growth investor backed and is currently operating.
When was Luminal founded?
Luminal was founded in 2025. It employs 11 to 50 people.
Where is Luminal based?
Luminal is headquartered in San Francisco, United States, in the North America region.
How does Luminal make money?
Two revenue lines are on record. Luminal Cloud - Serverless Inference is the primary driver. The others are on-Prem Deployment.
Who are Luminal's main competitors?
Direct peers on record are vLLM (project), Fireworks AI, Together AI, Anyscale, Modal, Replicate and Hugging Face Text Generation Inference (TGI). Broad incumbents are TensorRT-LLM (NVIDIA) and OctoAI (acquired by NVIDIA). DeepSpeed (Microsoft) is listed as an emerging player.
Does Luminal have an API?
No public API is recorded for Luminal.
What industry is Luminal in?
Luminal's product category is AI Inference Platform / ML Compiler Infrastructure. Its primary akta.pro industry code is HDAEANAC, Model Serving, Inference & Deployment Platforms (APIs, Edge/On-Prem), with a secondary code of HDAAABAF, Model Deployment, Serving & Inference Platforms. Its NAICS code is 513210 and its SIC code is 7372.