Wafer
Wafer is a seed-stage AI infrastructure company that delivers optimized open-source LLM inference via Serverless APIs and Dedicated Endpoints, using proprietary GPU kernels on AMD MI300X/MI355X and NVIDIA Blackwell to serve AI developers and compliance-bound enterprises.
- Company typePrivate
- Founded2025
- HeadquartersSan Francisco, United States
- Headcount1–10
- GTM typeB2B
- OfferingSoftware
What Wafer does
Wafer is a San Francisco-based private company founded in 2025 that provides optimized open-source large language model inference as a service, combined with a portfolio of GPU development tools for kernel engineers and AI coding agents. Its core technology centers on workload-specific GPU kernel engineering for AMD MI300X/MI355X and NVIDIA Blackwell hardware, including custom GEMM, MoE, and attention kernels, MXFP4/NVFP4 quantization, multi-token prediction speculative decoding, stepped-decode scheduling, and cache-aware routing that delivers claimed best-in-class throughput (310 tok/s on Qwen3.5 397B, 430 tok/s on Qwen3.6-35B-A3B, 11.33x speedup on Kimi 2.5) on Artificial Analysis benchmarks. The company monetizes through three primary streams: a self-serve Serverless API with pay-per-million-token pricing across six flagship open models, a Wafer Pass subscription with weekly/monthly/yearly tiers and included quotas, and quote-based Dedicated Endpoints provisioned in under 24 hours with HIPAA BAA, zero data retention, US-only data residency, and SLA-backed uptime for compliance-bound enterprise workloads. Distribution is hybrid: a product-led growth signup at app.wafer.ai for individual developers and a Cal.com-driven enterprise field-sales motion for Dedicated Endpoints, supplemented by strategic co-marketing partnerships with DigitalOcean and Parasail and a rapidly expanding dev-tools ecosystem (KernelArena, VS Code/Cursor extension, wafer-ai CLI, GPU Docs, Cloud Compiler Analyzer, Perfetto Trace Viewer, ROCprofiler Compute, Workspaces, Trace Compare, Chip Benchmark). The company has raised $4 million in seed funding led by Fifty Years with Y Combinator and Liquid 2 participation, operates with approximately six visible team members, and counts Jeff Dean, Woj Zaremba, Dan Fu, Charlie Songhurst, and Arash Ferdowsi among its advisors.
Wafer firmographics
Firmographics- Name
- Wafer
- Legal name
- Wafer, Inc.
- Website
- https://www.wafer.ai
- Company type
- Private
- Founded year
- 2025
- Operating status
- Operating
- Headcount range
- 1–10 employees
- Short description
- Wafer is a seed-stage AI infrastructure company that delivers optimized open-source LLM inference via Serverless APIs and Dedicated Endpoints, using proprietary GPU kernels on AMD MI300X/MI355X and NVIDIA Blackwell to serve AI developers and compliance-bound enterprises.
- Ownership category
- akta.pro rank
Wafer industry classification
Industry- Product category
- LLM Inference Platform
- NAICS
- Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (5182), Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (518)
- SIC
- Services-Prepackaged Software (7372)
- akta.pro primary industry
- Model Hosting, Serving & Inference Platforms (HDAAACAB)
- akta.pro secondary industry
- Model Deployment, Serving & Inference Platforms (HDAAABAF)
Keywords
Where Wafer is headquartered
LocationHeadquarters
- HQ city
- San Francisco
- HQ country
- United States
- HQ region
- North America
Offices1 record
Markets served
Wafer business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Technology or R&D, Personnel, Infrastructure, Operations, Marketing or Sales
Revenue model
- Serverless Inference (Pay-per-token): Pay-as-you-go serverless inference for open-source LLMs. Per-million-token pricing for input, output, and cache hits. Open models including GLM-5.1, GLM-5.2, Kimi-K2.6, Qwen 3.5, DeepSeek V4 Pro/Flash.
- Wafer Pass Subscription: Subscription plans (weekly, monthly, yearly) with included quotas and supported models. Renews automatically. Tiered pricing with overage rates for usage beyond included quota.
- Dedicated Endpoints: Dedicated inference infrastructure for mission-critical workloads. Custom-tuned deployments in under 24 hours, SLA-backed uptime (99.9% standard), TTFT-based SLA options, HIPAA BAA available.
- GPU Workspaces: On-demand GPU compute for coding agents. Pay-per-second only when commands are running, not for idle time. B200 and MI300X available.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Usage-based | Pay-as-you-go | Serverless inference for GLM-5.2 - flagships strongest coding/reasoning |
| Usage-based | Pay-as-you-go | Serverless inference for GLM-5.1 - flagship with strong coding/reasoning |
| Usage-based | Pay-as-you-go | Serverless inference for Kimi-K2.6 - sparse MoE with 262K context |
| Usage-based | Pay-as-you-go | Serverless inference for Qwen 3.5 397B-A17B - MoE flagship model |
| Usage-based | Pay-as-you-go | Serverless inference for DeepSeek V4 Pro - higher capability version |
| Usage-based | Pay-as-you-go | Serverless inference for DeepSeek V4 Flash - fast, cost-efficient version |
| Other | Multi-year contract | Dedicated endpoints for sensitive/mission-critical workloads |
| Subscription | Multi-year contract | Wafer Pass subscription plans |
Go-to-market motion3 records
Distribution channels4 records
Marketing channels6 records
Wafer product offering
Product offeringCore offering
Wafer provides serverless and dedicated inference APIs for open-source large language models (GLM-5.2, GLM-5.1, Kimi-K2.6, Qwen 3.5, DeepSeek V4 Pro/Flash) optimized through proprietary custom GPU kernels on AMD MI300X/MI355X and NVIDIA Blackwell hardware. The platform offers an OpenAI-compatible Chat Completions API with workload-specific tuning, zero data retention options, and dedicated endpoints with SLA-backed uptime and HIPAA BAA for compliance-bound enterprise workloads.
Product overview
Wafer provides fast open-source LLM inference for enterprise through two core offerings: Wafer Serverless (pay-as-you-go API access to open models) and Wafer Dedicated Inference (customized dedicated endpoints with SLA-backed reliability). The Serverless platform serves models including GLM-5.2, GLM-5.1, Kimi-K2.6, Qwen 3.5, DeepSeek V4 Pro, and DeepSeek V4 Flash via an OpenAI-compatible API. Wafer also develops a suite of GPU development tools including the wafer-ai CLI, VS Code/Cursor extension, GPU Docs, Cloud Compiler Analyzer, Perfetto Trace Viewer, ROCprofiler Compute, Workspaces, Chip Benchmark, KernelArena, and Trace Compare. The company has released custom model optimizations including wafer-ai/Kimi-K2.6-NVFP4 for Blackwell and has achieved leading inference performance on AMD MI355X hardware through custom kernel engineering.
Differentiator
Problem solved
Functional benefit
Products and services
- Wafer Serverless
Quantifiable outcome
- 30% reduction in voice-agent TTFT (800ms to 550ms) at 25% higher peak load
- +4 more outcomes
Companies that use Wafer
Customer profileNamed customers3 records
Segments3 records
Ideal customer profiles1 record
Wafer technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration6 records
AI capability3 records
Feature12 records
Wafer partnerships and signals
Strategic signalPartnerships
Eleven partnerships are on record, tiered key advisor, technical advisor, key partner, advisor and infrastructure partner.
- Jeff Dean (Google Chief Scientist)key advisorChief Scientist at Google serving as advisor to Wafer. Provides deep technical guidance on large-scale systems and AI infrastructure.
- Woj Zaremba (OpenAI Co-founder)key advisorCo-founder of OpenAI serving as advisor. Provides perspective on frontier model development and AI safety.
- Dan Fu (Together Head of Kernels)technical advisorHead of Kernels at Together serving as advisor. Brings expertise in kernel optimization for AI inference.
- DigitalOceankey partnerJoint case study demonstrating 11.3x speedup on Kimi 2.5, 7.32x on DeepSeek V3.2 using DigitalOcean's AMD GPU infrastructure. DigitalOcean blog cross-post. Strategic partnership for AMD inference optimization.
- Parasailkey partnerPartnership to launch wafer-ai/Kimi-K2.6-NVFP4 - Blackwell NVFP4 quantized weights for production Kimi K2.6 inference. Achieved up to 58% more throughput vs INT4, 43% faster token streaming.
- Charlie Songhurst (Meta Board)advisorBoard of Directors at Meta serving as advisor. Provides perspective on large-scale AI deployments and business strategy.
- Arash Ferdowsi (Dropbox Co-founder)advisorCo-founder of Dropbox serving as advisor. Brings experience building and scaling successful tech companies.
- TensorWaveinfrastructure partnerInfrastructure sub-processor for Covered Processing under DPA. AMD GPU cloud provider. Named as one of Wafer's infrastructure partners.
- Vultrinfrastructure partnerInfrastructure sub-processor for Covered Processing under DPA. Cloud GPU provider.
- AWSinfrastructure partnerInfrastructure sub-processor (Amazon Web Services, Inc.) for Covered Processing. GPU cloud services.
- Nebiusinfrastructure partnerInfrastructure sub-processor (Nebius B.V.) for Covered Processing under DPA. Alternative cloud GPU provider.
Scale indicators13 records
Recent moves6 records
Expansion highlights5 records
Wafer competitors and assessment
Company assessmentMarket position
Competitive moat5 records
Key risks7 records
Key highlights8 records
Customer concentration
Wafer social profiles
Digital presenceWafer compliance and trust
Trust signalCompliance2 records
Wafer financial estimates
Financial estimateRevenue estimate
Valuation estimate
Wafer leadership team
Management profileNumber of profiles
Profiles6 records
Wafer funding detail
Funding detailFunding overview
Funding rounds3 records
Investors9 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Wafer M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Wafer
What does Wafer do?
Wafer provides serverless and dedicated inference APIs for open-source large language models (GLM-5.2, GLM-5.1, Kimi-K2.6, Qwen 3.5, DeepSeek V4 Pro/Flash) optimized through proprietary custom GPU kernels on AMD MI300X/MI355X and NVIDIA Blackwell hardware. The platform offers an OpenAI-compatible Chat Completions API with workload-specific tuning, zero data retention options, and dedicated endpoints with SLA-backed uptime and HIPAA BAA for compliance-bound enterprise workloads.
Is Wafer a public or private company?
Wafer is a private company. It is classified as venture growth investor backed and is currently operating.
When was Wafer founded?
Wafer was founded in 2025. It employs 1 to 10 people.
Where is Wafer based?
Wafer is headquartered in San Francisco, United States, in the North America region.
How does Wafer make money?
Four revenue lines are on record. Serverless Inference (Pay-per-token) is the primary driver. The others are wafer Pass Subscription, dedicated Endpoints and GPU Workspaces.
Does Wafer have an API?
Yes. Wafer offers an inference API that follows the OpenAI Chat Completions schema, enabling drop-in replacement for OpenAI API. Supports streaming, tool use (function calling), and JSON mode across every Serverless model. Existing clients including the OpenAI SDK, LangChain, LiteLLM, and agent harnesses like Claude Code or Cline work by swapping the base URL and API key. Available as Serverless (pay-as-you-go) and Dedicated endpoints for enterprise workloads. Developer documentation is at app.wafer.ai/signup.
What industry is Wafer in?
Wafer's product category is LLM Inference Platform. Its primary akta.pro industry code is HDAAACAB, Model Hosting, Serving & Inference Platforms, with a secondary code of HDAAABAF, Model Deployment, Serving & Inference Platforms. Its NAICS code is 5182 and its SIC code is 7372.