Lilac
Lilac is a two-sided marketplace that routes AI inference onto idle enterprise GPUs via a Kubernetes operator, offering pay-per-token access to frontier open-weight models (Kimi K2.6, GLM 5.1/5.2, Gemma 4, MiniMax M2.7/M3) through an OpenAI-compatible API, with subscriptions, batch compute, and cluster reservations for AI developers and small teams.
- Company typePrivate
- Founded2025
- HeadquartersSan Francisco, United States
- Headcount11–50
- GTM typeB2B
- OfferingSoftware
What Lilac does
Lilac Research Inc. is a San Francisco-based two-sided marketplace for idle enterprise GPU capacity, founded in 2025 and backed by Y Combinator. On the supply side, GPU cluster operators install Lilac's Kubernetes-based GPU Operator (a controller with a 30-second reconciliation loop and LIFO preemption that completes in under 60 seconds) to monetize idle NVIDIA capacity that would otherwise sit unused; Lilac routes inference traffic into that capacity and shares 70% of gross token revenue with suppliers, retaining 30% as a marketplace fee. On the demand side, AI developers and engineering teams access frontier open-weight models — Kimi K2.6 (Moonshot AI), GLM 5.1 and 5.2 (Z.ai), Gemma 4 (Google), and MiniMax M2.7 and M3 — through an OpenAI-compatible REST API at api.getlilac.com, with pay-per-token usage-based pricing, no contracts or minimums, cache read discounts for repeated context, and subscription plans (Basic $10/mo, Pro $30/mo, Max $100/mo) that layer in live supply-linked discounts of up to 75%.
Beyond chat inference, Lilac sells per-second batch container compute on idle H100 ($1.00/hr) and H200 ($1.50/hr) capacity and brokers committed cluster reservations (H100 to B300) from neo-cloud partners, with enterprise customers reached via founder-led direct sales and monthly invoicing. The go-to-market is product-led and self-serve, accelerated by YC distribution, a Discord community, technical blog content on idle-GPU economics, and drop-in integration with OpenAI-compatible developer tools including OpenCode, Cursor, and Continue. As of early 2026 the company has roughly $3M in disclosed funding ($500K YC in May 2025 and a $2.5M round in October 2025) and operates with 11-50 employees, serving a long tail of AI startups (Osmosis, Understudy Labs, NanoGPT, Eden AI, Linum, Infron, Coral Bricks, Floot, Channel3) alongside model-provider partners and select enterprise teams.
Lilac firmographics
Firmographics- Name
- Lilac
- Legal name
- Lilac Research Inc.
- Website
- https://getlilac.com
- Company type
- Private
- Founded year
- 2025
- Operating status
- Operating
- Headcount range
- 11–50 employees
- Short description
- Lilac is a two-sided marketplace that routes AI inference onto idle enterprise GPUs via a Kubernetes operator, offering pay-per-token access to frontier open-weight models (Kimi K2.6, GLM 5.1/5.2, Gemma 4, MiniMax M2.7/M3) through an OpenAI-compatible API, with subscriptions, batch compute, and cluster reservations for AI developers and small teams.
- Ownership category
- akta.pro rank
Lilac industry classification
Industry- Product category
- AI Inference Cloud
- akta.pro primary industry
- Decentralized GPU/AI Compute Networks (training/inference, GPU renting, model hosting) (FSAPALAC)
- akta.pro secondary industry
- AI Compute Cloud & GPU-as-a-Service (HDAAAAAK)
Keywords
Where Lilac is headquartered
LocationHeadquarters
- HQ city
- San Francisco
- HQ country
- United States
- HQ region
- North America
Offices1 record
Markets served
Lilac business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Technology or R&D, Personnel, Infrastructure, Operations, Marketing or Sales
Revenue model
- Per-token Inference Revenue: Lilac charges customers per token for inference on open frontier models (Kimi K2.6, GLM 5.1, GLM 5.2, Gemma 4, MiniMax M2.7, MiniMax M3) routed through idle enterprise GPUs. Pricing is per million tokens with no contracts or minimums. This is the primary revenue stream from AI developers and teams using the API.
- Subscription Plans (Personal): Monthly subscription plans (Basic $10/mo, Pro $30/mo, Max $100/mo) bundle dollar-denominated included model usage with live per-model discounts. Included usage does not roll over. Subscriptions are for individual, non-commercial use only. Commercial workloads must use prepaid credits or enterprise billing.
- Batch / Container Jobs: Batch GPU compute jobs billed per second at $1.00/hr for H100s and $1.50/hr for H200s. Customers submit a container image and command; Lilac runs on spare GPU capacity. Used for training, fine-tuning, and batch processing workloads.
- Supplier Revenue Share (30%): Lilac takes 30% of gross token revenue from inference processed on supplier GPU infrastructure. Suppliers keep 70%. Payouts are monthly via wire transfer or ACH. This is Lilac's take rate from the two-sided marketplace.
- Cluster Reservations: For enterprise customers needing dedicated GPU capacity (H100, H200, B200, B300), Lilac brokers competitive quotes from neo-cloud partners and helps offset cost with idle-GPU software. Estimated pricing ~$2.00/hr for H100 on 1-month reservation.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Usage-based | Pay-as-you-go | Kimi K2.6 - $0.70/M input, $0.20/M cache read, $3.50/M output |
| Usage-based | Pay-as-you-go | GLM 5.1 - $0.90/M input, $0.27/M cache read, $3.00/M output |
| Usage-based | Pay-as-you-go | GLM 5.2 - $0.90/M input, $0.27/M cache read, $3.00/M output |
| Usage-based | Pay-as-you-go | Gemma 4 (31B) - $0.11/M input, $0.35/M output |
| Usage-based | Pay-as-you-go | MiniMax M2.7 - $0.30/M input, $0.055/M cache read, $1.20/M output |
| Usage-based | Pay-as-you-go | MiniMax M3 - $0.28/M input, $0.05/M cache read, $1.10/M output |
| Subscription | Monthly | Basic - $10/mo, up to $80 value (2x multiplier) |
| Subscription | Monthly | Pro - $30/mo, up to $300 value (2.5x multiplier) |
| Subscription | Monthly | Max - $100/mo, up to $1200 value (3x multiplier) |
| Usage-based | Pay-as-you-go | Batch H100 - $1.00/hr |
| Usage-based | Pay-as-you-go | Batch H200 - $1.50/hr |
| Usage-based | Multi-year contract | Cluster H100 Reservation - ~$2.00/hr (1 month estimate) |
Go-to-market motion2 records
Distribution channels4 records
Marketing channels6 records
Lilac product offering
Product offeringCore offering
Lilac runs a two-sided idle GPU network: customers access pay-per-token inference on frontier open-weight models (Kimi K2.6, GLM 5.1, GLM 5.2, Gemma 4, MiniMax M2.7, MiniMax M3) via an OpenAI-compatible API, plus subscription plans, container batch compute, and brokered GPU cluster reservations. On the supply side, GPU operators install Lilac's Kubernetes operator to monetize idle GPU capacity, keeping 70% of token revenue while Lilac takes a 30% marketplace commission.
Product overview
Lilac is a unified idle GPU network platform offering four primary compute services: Inference (pay-per-token GPU inference on idle enterprise GPUs via OpenAI-compatible API), Subscriptions (monthly plans with up to 12x value through live discounts), Batch (container job processing on spare GPU capacity), and Clusters (dedicated GPU reservation with idle-GPU cost offset). The platform operates a two-sided market: customers access affordable frontier AI model inference while GPU operators monetize idle capacity through the Lilac GPU Operator (a Kubernetes controller) and GPU Pool configuration system. Lilac supports multiple frontier models including Kimi K2.6, GLM 5.1, GLM 5.2, Gemma 4, MiniMax M2.7, and MiniMax M3, accessible via OpenAI-compatible API endpoints.
Differentiator
Problem solved
Functional benefit
Products and services
- Inference (GPU Inference API) Pay-per-token GPU inference on idle enterprise GPUs. Runs frontier open-weight models (Kimi K2.6, GLM 5.1, GLM 5.2, Gemma 4, MiniMax M2.7, MiniMax M3) via an OpenAI-compatible API with no contracts or minimums.
- Subscriptions Monthly subscription plans (Basic $10/mo, Pro $30/mo, Max $100/mo) that bundle dollar-denominated model usage with live per-model discounts up to 75% off, delivering up to 12x value vs prepaid credits. Individual, non-commercial use.
- Batch Container job processing on spare GPU capacity. Customers submit a container image and command; Lilac runs it on spot-style idle GPU compute at $1/hr H100s and $1.50/hr H200s, priced per second. Used for training, fine-tuning, and batch processing.
- Clusters GPU cluster reservation service. Lilac sources competitive quotes from neo-cloud partners for dedicated GPU clusters (H100, H200, B200, B300) and helps offset cost with idle-GPU software. ~$2.00/hr for H100 on 1-month reservation.
- GPU Supplier Network Network for GPU operators to monetize idle GPU capacity. Suppliers install the Lilac Kubernetes operator to automatically pick up inference traffic when GPUs are idle, earning 70% of gross token revenue (Lilac keeps 30%).
- Lilac GPU Operator Kubernetes controller deployed on supplier clusters (Helm chart, v0.3.5). Discovers idle GPUs, communicates with Lilac's control plane, and manages inference workload pods using a customized vLLM fork. Runs a 30-second reconciliation sync loop for capacity calculation and workload scheduling, with LIFO preemption and sub-60-second graceful shutdown.
Quantifiable outcome
- 30% cheaper tokens vs typical on-demand GPU inference rates
- +5 more outcomes
Companies that use Lilac
Customer profileNamed customers13 records
Segments4 records
Ideal customer profiles4 records
Lilac technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration3 records
AI capability7 records
Feature6 records
Lilac partnerships and signals
Strategic signalPartnerships
Six partnerships are on record, tiered flagship and core.
- MiniMaxflagshipLilac partnered with MiniMax to bring commercially licensed MiniMax M2.7 to Lilac's network. MiniMax provides open-weight models with commercial licensing; Lilac handles the hosted inference endpoint and routing onto GPU capacity. The partnership gives customers a clear commercial path to use the model without managing deployments.
- Moonshot AIcoreMoonshot AI's Kimi K2.6 model (1T total parameters, 32B activated, MoE architecture) is served on Lilac's idle GPU network. Kimi K2.6 supports 262K context, multimodal input, and reasoning. Lilac routes inference traffic to capacity providers running the model.
- Z.ai (Zhipu AI)coreZ.ai's GLM 5.1 (754B parameters, 202.8K context) and GLM 5.2 (524K context) are served on Lilac's idle GPU network. GLM 5.1 is MIT licensed. GLM 5.2 supports configurable reasoning effort and tool calling for agentic workflows.
- AWS (Amazon Web Services)coreAWS is used as cloud infrastructure for Lilac's operator Helm chart (hosted on AWS ECR public registry). Capacity Providers may run on AWS infrastructure. Subprocessor for cloud hosting.
- vLLMcoreLilac's inference pods run a customized fork of vLLM, a high-performance inference engine. vLLM handles model serving, batching, and GPU memory management. Lilac's fork is tuned for idle-GPU scheduling and shared warm endpoints.
- StripecoreStripe handles all payment processing for Lilac's credit-based billing system. Subscriptions, auto top-up, and enterprise invoicing all run through Stripe. Listed as subprocessors for payment processing.
Scale indicators8 records
Recent moves6 records
Expansion highlights6 records
Lilac competitors and assessment
Company assessmentDirect peers
- Together AI: Together AI provides a GPU cloud and inference API for open-source and open-weight AI models, serving frontier LLMs with per-token pricing. Direct competitor in the AI inference cloud market with overlapping customer base of AI developers and engineering teams.
- Fireworks AI: Fireworks AI offers a low-latency inference cloud for open-weight LLMs and custom models, with per-token pricing and an OpenAI-compatible API. Closely comparable business model to Lilac's inference product targeting similar developer and enterprise customers.
- Vast.ai: Vast.ai operates a decentralized marketplace where GPU operators rent idle capacity to AI customers. Most structurally similar peer to Lilac—same two-sided idle-GPU marketplace model, though Vast.ai is more consumer/developer oriented and less focused on frontier LLM inference.
- RunPod: RunPod provides GPU cloud services including on-demand and spot GPU rentals for AI training and inference. Overlaps with Lilac's batch and inference products, serving similar AI developer and researcher customers.
- Anyscale: Anyscale offers a compute platform for AI workloads built on Ray, including managed inference and training on GPU clusters. Competes with Lilac for enterprise AI team workloads and provides similar OpenAI-compatible inference endpoints.
- Modal Labs: Modal provides serverless compute for AI workloads, including GPU-backed batch jobs and inference endpoints. Direct overlap with Lilac's batch and inference products with developer-first PLG motion.
- Replicate: Replicate runs an API platform for hosting and serving open-source AI models on GPU infrastructure, with pay-per-prediction billing. Closely comparable developer-first inference cloud targeting similar AI application builders.
Broad incumbents
- CoreWeave: CoreWeave is a large-scale GPU cloud provider specializing in AI workloads with dedicated NVIDIA GPU capacity. Competes with Lilac's cluster reservation offering and broader inference market as a well-capitalized incumbent with completed enterprise certifications.
- Lambda Labs: Lambda provides GPU cloud services and AI inference APIs, with on-demand and reserved capacity for training and inference. Larger, more established peer competing in the same GPU inference and cluster markets as Lilac.
Emerging players
- io.net: io.net operates a decentralized GPU network for AI compute, aggregating idle GPUs from data centers and crypto miners to serve AI training and inference workloads. Comparable to Lilac's idle GPU aggregation model with a blockchain/crypto-economic layer.
Market position
Strengths5 records
Weaknesses4 records
Competitive moat5 records
Key risks5 records
Key highlights6 records
Customer concentration
Lilac social profiles
Digital presenceLilac compliance and trust
Trust signalCompliance1 record
Lilac financial estimates
Financial estimateRevenue estimate
Valuation estimate
Lilac leadership team
Management profileNumber of profiles
Profiles2 records
Lilac funding detail
Funding detailFunding overview
Funding rounds3 records
Investors3 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Lilac M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Lilac
What does Lilac do?
Lilac runs a two-sided idle GPU network: customers access pay-per-token inference on frontier open-weight models (Kimi K2.6, GLM 5.1, GLM 5.2, Gemma 4, MiniMax M2.7, MiniMax M3) via an OpenAI-compatible API, plus subscription plans, container batch compute, and brokered GPU cluster reservations. On the supply side, GPU operators install Lilac's Kubernetes operator to monetize idle GPU capacity, keeping 70% of token revenue while Lilac takes a 30% marketplace commission.
Is Lilac a public or private company?
Lilac is a private company. It is classified as venture growth investor backed and is currently operating.
When was Lilac founded?
Lilac was founded in 2025. It employs 11 to 50 people.
Where is Lilac based?
Lilac is headquartered in San Francisco, United States, in the North America region.
How does Lilac make money?
Five revenue lines are on record. Per-token Inference Revenue is the primary driver. The others are subscription Plans (Personal), batch / Container Jobs, supplier Revenue Share (30%) and cluster Reservations.
Who are Lilac's main competitors?
Direct peers on record are Together AI, Fireworks AI, Vast.ai, RunPod, Anyscale, Modal Labs and Replicate. Broad incumbents are CoreWeave and Lambda Labs. io.net is listed as an emerging player.
Does Lilac have an API?
Yes. Lilac provides a real-time OpenAI-compatible inference API for GPU inference. The API is public and supports per-token pricing with no minimums or contracts. Key endpoints include POST /v1/chat/completions, POST /v1/completions (legacy), POST /v1/responses, and GET /v1/models. Authentication is via Bearer token (API key). Default rate limit is 200 requests per minute per organization; higher limits available upon request. The API supports streaming, tool/function calling, structured output (JSON mode), vision (image inputs), and reasoning controls. Available at https://api.getlilac.com/v1. Developer documentation is at docs.getlilac.com.
What industry is Lilac in?
Lilac's product category is AI Inference Cloud. Its primary akta.pro industry code is FSAPALAC, Decentralized GPU/AI Compute Networks (training/inference, GPU renting, model hosting), with a secondary code of HDAAAAAK, AI Compute Cloud & GPU-as-a-Service.