Gimlet Labs
Gimlet Labs builds a multi-silicon serverless inference cloud for AI agents, orchestrating workloads across CPUs, GPUs, and SRAM-centric accelerators to deliver 3-10x performance gains, primarily serving frontier AI labs and hyperscalers.
- Company typePrivate
- Founded2023
- HeadquartersSan Francisco, United States
- Headcount11–50
- GTM typeB2B
- OfferingSoftware
What Gimlet Labs does
Gimlet Labs is a San Francisco-based applied AI research and infrastructure company founded in 2023 that builds software for executing agent-based AI workloads across heterogeneous hardware. Its core product, Gimlet Cloud, is a serverless inference platform that orchestrates AI workloads across CPUs, NVIDIA GPUs, Intel Gaudi, AMD accelerators, ARM, Cerebras, and d-Matrix SRAM-centric accelerators, automatically mapping each phase of inference to the optimal processor to deliver 3-10x performance gains at the same cost and power. The company also ships kforge, an autonomous kernel-generation toolkit that uses a multi-agent system with shared memory to produce optimized low-level kernels directly from PyTorch across CUDA, ROCm, and Metal backends, with disclosed speedups of ~1.8x on H100 and ~1.24x on Apple M4 over PyTorch baselines.
The platform is positioned for agentic AI workloads that produce 5-15x more tokens than chat models, addressing an inference bottleneck that the company says leaves current data centers operating at only 15-30% utilization. Revenue is generated through consumption-based pricing on Gimlet Cloud inference compute, augmented by a standalone kforge offering and supplemented by enterprise and self-serve channels. Go-to-market operates on a dual track: high-touch enterprise field sales targeting frontier AI labs and hyperscalers, alongside product-led growth through the hosted platform for developers and smaller teams.
As of March 2026, Gimlet has raised $92 million in total funding ($80M Series A led by Menlo Ventures at a $3.1B post-money valuation plus $12M seed), disclosed eight-figure revenue exceeding $10 million within five months of its October 2025 stealth launch, and tripled its customer base. The company operates with 11-50 employees from 255 Potrero Ave in San Francisco, and has established technology or strategic relationships with NVIDIA, Intel, AMD, ARM, Cerebras, d-Matrix, MLCommons, and a deep advisor bench spanning Stanford faculty and operators from Sequoia, Figma, VMware, and SiMA.ai.
Gimlet Labs firmographics
Firmographics- Name
- Gimlet Labs
- Legal name
- Gimlet Labs, Inc.
- Website
- https://gimletlabs.ai
- Company type
- Private
- Founded year
- 2023
- Operating status
- Operating
- Headcount range
- 11–50 employees
- Short description
- Gimlet Labs builds a multi-silicon serverless inference cloud for AI agents, orchestrating workloads across CPUs, GPUs, and SRAM-centric accelerators to deliver 3-10x performance gains, primarily serving frontier AI labs and hyperscalers.
- Ownership category
- akta.pro rank
Gimlet Labs industry classification
Industry- Product category
- AI Inference Infrastructure
- NAICS
- Computer Systems Design Services (541512)
- SIC
- Services-Computer Programming, Data Processing, Etc. (7370)
- akta.pro primary industry
- Model Deployment, Serving & Inference Platforms (HDAAABAF)
- akta.pro secondary industries
- AI Compute Virtualization & Scheduling (GPU virtualization, cluster schedulers) (HDAAAAAG), On-Device Inference Runtimes & SDKs (mobile/embedded) (HDAAAJAB)
Keywords
Where Gimlet Labs is headquartered
LocationHeadquarters
- HQ city
- San Francisco
- HQ country
- United States
- HQ region
- North America
Offices1 record
Markets served
Gimlet Labs business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Technology or R&D, Infrastructure, Personnel, Operations, Marketing or Sales
Revenue model
- Serverless AI Inference Platform (Gimlet Cloud): Gimlet Cloud provides serverless inference for AI agents, handling scheduling, orchestration, and optimization across heterogeneous hardware. Users can run simple agents to complex multi-agent systems with custom logic and data sources. Revenue is generated through consumption-based pricing for inference compute used.
- Autonomous Kernel Generation Tool (kforge): kforge autonomously generates optimized low-level kernels directly from PyTorch code. May be offered as a standalone commercial product or integrated offering.
Go-to-market motion2 records
Distribution channels3 records
Marketing channels5 records
Gimlet Labs product offering
Product offeringCore offering
Gimlet Labs builds and operates Gimlet Cloud, a serverless multi-silicon AI inference platform that runs agentic workloads (simple agents through multi-agent systems with custom logic and data sources) by automatically disaggregating and orchestrating each phase (prefill, decode, speculative decoding, tool calls) across heterogeneous hardware including NVIDIA GPUs, Intel Gaudi, AMD, ARM, Cerebras, and d-Matrix Corsair accelerators. The company also offers kforge, an autonomous kernel generation toolkit that produces optimized low-level kernels directly from PyTorch code across CUDA, ROCm, and Metal backends.
Product overview
Gimlet Labs is an applied AI research and product company that offers two core products: Gimlet Cloud (a serverless inference platform for AI agents) and kforge (an autonomous kernel generation toolkit). Gimlet Cloud provides multi-silicon inference capabilities, distributing AI workloads across CPUs, GPUs, and specialized accelerators from vendors including NVIDIA, AMD, Intel, ARM, Cerebras, and d-Matrix. kforge complements the platform by autonomously generating optimized low-level kernels for PyTorch across diverse hardware. Together, these products aim to make AI workloads 10X more efficient by expanding the pool of usable compute and improving how it is orchestrated.
Differentiator
Problem solved
Functional benefit
Brands
- Gimlet Cloud: Serverless inference cloud platform for AI agents. Enables running everything from simple agents to complex multi-agent systems with custom logic and data sources. The platform handles scheduling, orchestration, and optimization across heterogeneous hardware including CPUs, GPUs, and specialized AI accelerators.
- kforge
Products and services
- Gimlet Cloud Serverless inference cloud platform for AI agents that enables running simple agents to complex multi-agent systems with custom logic and data sources. The platform handles scheduling, orchestration, and optimization across heterogeneous hardware including CPUs, GPUs, and specialized AI accelerators from NVIDIA, AMD, Intel, ARM, Cerebras, and d-Matrix.
- kforge Autonomous kernel generation toolkit that generates optimized low-level kernels directly from PyTorch code using a multi-agent system with shared memory to explore designs, enforce correctness checks, and identify the fastest kernels across CUDA, ROCm, and Metal backends.
Quantifiable outcome
- 3-10x speedup on >1T parameter frontier models with large context windows at same cost and power
- +5 more outcomes
Companies that use Gimlet Labs
Customer profileNamed customers2 records
Segments3 records
Ideal customer profiles3 records
Gimlet Labs technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- No
- API docs
- API detail
Core technology
AI maturity
App detail
AI capability8 records
Feature9 records
Gimlet Labs partnerships and signals
Strategic signalPartnerships
Twelve partnerships are on record, tiered minor and core.
- MLCommonsminorGimlet Labs joined MLCommons as a member company to contribute to new benchmarks for agentic inference and support open standards for ML performance across industry and research. This positions Gimlet to help shape emerging performance standards for heterogeneous inference systems as agentic inference becomes a dominant workload.
- d-Matrixcored-Matrix and Gimlet Labs partnered to integrate d-Matrix Corsair SRAM-centric accelerators into Gimlet Cloud for AI inference workloads, achieving up to 10x improvements in latency and throughput per watt compared to GPU-only deployments. The partnership routes memory-sensitive operations like speculative decoding and token generation to Corsair while GPUs handle compute-heavy inference stages, delivering the best of both architectures within Gimlet's multi-silicon orchestration.
- NVIDIAcoreGimlet has established partnerships with NVIDIA to ensure compatibility of its multi-silicon inference cloud with NVIDIA hardware. NVIDIA GPUs serve as primary compute accelerators within the heterogeneous stack, with Gimlet optimizing workloads including prefill and verification stages on NVIDIA B200 and H100 GPUs.
- IntelcoreIntel Gaudi 3 accelerators are integrated into Gimlet's heterogeneous inference stack. Research demonstrates that B200:Gaudi3 multi-vendor disaggregation (NVIDIA B200 for prefill, Intel Gaudi 3 for decode) achieves 1.7-4x TCO improvements versus single-vendor disaggregation.
- Mark HorowitzminorMark Horowitz, Rambus Founder and Stanford Faculty, serves as an advisor to Gimlet Labs, contributing expertise in chip architecture and memory systems.
- Chris RéminorChris Ré, Entrepreneur and Stanford Faculty, serves as an advisor, contributing deep expertise in machine learning systems and scalable AI infrastructure.
- Lukas BiewaldminorLukas Biewald, Co-Founder of Weights & Biases, serves as an advisor, contributing expertise in ML tooling and experiment tracking infrastructure.
- Dylan FieldminorDylan Field, Founder and CEO of Figma, serves as an advisor, contributing expertise in product development and scaling consumer/enterprise software.
- Amarjit GillminorAmarjit Gill, Founder of P.A. Semi, serves as an advisor, contributing expertise in low-power processor design and chip company building.
- Krishna RangasayeeminorKrishna Rangasayee, Founder and CEO of SiMA.ai, serves as an advisor, contributing expertise in AI chip design and edge AI deployment.
- Andrew TanminorAndrew Tan, Founder of TLDR AI, serves as an advisor, contributing expertise in AI summarization products and developer-facing AI tools.
- Rangarajan RaghuramanminorRangarajan Raghuraman, former CEO of VMware, serves as an advisor, contributing expertise in enterprise infrastructure, virtualization, and large-scale distributed systems.
Scale indicators6 records
Recent moves7 records
Expansion highlights6 records
Gimlet Labs competitors and assessment
Company assessmentDirect peers
- Anyscale: Anyscale provides the Ray-based Anyscale Platform for distributed AI compute, including production LLM serving and agentic workloads. It competes with Gimlet on serverless AI inference and multi-node orchestration for production deployments.
- Modal Labs: Modal provides serverless cloud infrastructure for running AI workloads, including model inference and agentic pipelines. It targets a similar developer-led and enterprise adoption curve as Gimlet Cloud.
- Together AI: Together AI operates a cloud platform for training, fine-tuning, and serving open and custom AI models, including agentic inference workloads. It directly competes with Gimlet Cloud in the inference-as-a-service layer targeting frontier labs and AI developers.
- Replicate: Replicate runs a serverless cloud for running open-source ML models via API, including LLMs and multimodal agents. It competes with Gimlet Cloud on consumption-based inference pricing for AI applications.
- OctoAI (NVIDIA): OctoAI built a serverless inference platform for generative AI models and was acquired by NVIDIA. It is a direct predecessor-style competitor to Gimlet Cloud, with NVIDIA now able to bundle similar capabilities into its DGX Cloud stack.
- Fireworks AI: Fireworks AI offers a serverless inference platform optimized for low-latency LLM and agent inference, with fine-tuning and custom deployment options. It competes head-to-head with Gimlet Cloud on inference performance and developer experience.
Broad incumbents
- Hugging Face: Hugging Face offers Inference Endpoints and a broader AI platform spanning model hosting, training, and deployment. It competes with Gimlet for developer attention on hosted inference, though it spans a much broader model and tooling portfolio.
- CoreWeave: CoreWeave is a large GPU-cloud provider offering AI training and inference infrastructure at scale. While not specialized for agentic inference, it competes for the same hyperscaler and frontier-lab workloads and provides an alternative path for heterogeneous compute.
- Lambda Labs: Lambda is a GPU cloud and on-prem AI infrastructure provider serving labs and enterprises with training and inference clusters. It overlaps with Gimlet on serving AI workloads on GPU hardware, though without Gimlet's multi-silicon orchestration layer.
Emerging players
- RunPod: RunPod provides GPU cloud and serverless endpoints for AI inference and training. It targets developers and smaller teams seeking lower-cost alternatives to hyperscaler GPU clouds, partially overlapping with Gimlet's self-serve/PLG motion.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat4 records
Key risks6 records
Key highlights7 records
Customer concentration
Gimlet Labs social profiles
Digital presenceGimlet Labs financial estimates
Financial estimateRevenue estimate
Valuation estimate
Gimlet Labs leadership team
Management profileNumber of profiles
Profiles17 records
Gimlet Labs funding detail
Funding detailFunding overview
Funding rounds3 records
Investors18 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Gimlet Labs M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Gimlet Labs
What does Gimlet Labs do?
Gimlet Labs builds and operates Gimlet Cloud, a serverless multi-silicon AI inference platform that runs agentic workloads (simple agents through multi-agent systems with custom logic and data sources) by automatically disaggregating and orchestrating each phase (prefill, decode, speculative decoding, tool calls) across heterogeneous hardware including NVIDIA GPUs, Intel Gaudi, AMD, ARM, Cerebras, and d-Matrix Corsair accelerators. The company also offers kforge, an autonomous kernel generation toolkit that produces optimized low-level kernels directly from PyTorch code across CUDA, ROCm, and Metal backends.
Is Gimlet Labs a public or private company?
Gimlet Labs is a private company. It is classified as venture growth investor backed and is currently operating.
When was Gimlet Labs founded?
Gimlet Labs was founded in 2023. It employs 11 to 50 people.
Where is Gimlet Labs based?
Gimlet Labs is headquartered in San Francisco, United States, in the North America region.
How does Gimlet Labs make money?
Two revenue lines are on record. Serverless AI Inference Platform (Gimlet Cloud) is the primary driver. The others are autonomous Kernel Generation Tool (kforge).
Who are Gimlet Labs's main competitors?
Direct peers on record are Anyscale, Modal Labs, Together AI, Replicate, OctoAI (NVIDIA) and Fireworks AI. Broad incumbents are Hugging Face, CoreWeave and Lambda Labs. RunPod is listed as an emerging player.
Does Gimlet Labs have an API?
No public API is recorded for Gimlet Labs.
What industry is Gimlet Labs in?
Gimlet Labs's product category is AI Inference Infrastructure. Its primary akta.pro industry code is HDAAABAF, Model Deployment, Serving & Inference Platforms, with a secondary code of HDAAAAAG, AI Compute Virtualization & Scheduling (GPU virtualization, cluster schedulers). Its NAICS code is 541512 and its SIC code is 7370.