FriendliAI
FriendliAI is a frontier AI inference cloud that deploys, scales, and monitors large language and multimodal models for AI startups, SaaS firms, and large enterprises. Its optimized GPU serving stack claims up to 3x faster inference and multi-cloud distribution via AWS, OCI, Samsung, and Nebius.
- Company typePrivate
- Founded2021
- HeadquartersRedwood City, United States
- Headcount11–50
- GTM typeB2B
- OfferingSoftware
What FriendliAI does
FriendliAI operates a frontier AI inference cloud that enables enterprises and AI-native companies to deploy, scale, and monitor large language, vision, and multimodal models in production. Founded in 2021 by researchers from Seoul National University who pioneered continuous batching, the company markets a purpose-built GPU serving stack that combines model-level optimizations (custom GPU kernels, speculative decoding, Host KV Cache, parallel inference) with infrastructure-level optimizations (advanced caching, geo-distributed multi-cloud scaling) to deliver inference that is claimed to be up to 3x faster than vLLM and 50%-90% cheaper than closed-model APIs. The platform supports day-0 availability for frontier open-weight models including GLM-5.2, DeepSeek V4, Kimi K2.6, NVIDIA Nemotron 3, Llama-3.3-70B, Qwen3, and Gemma-4-31B-it, and offers Anthropic Messages API compatibility for drop-in client integration.
The product portfolio is organized around the Friendli Suite console and comprises three core commercial offerings: Model APIs (serverless, pay-per-token inference for frontier open-weight models across text, vision, and speech-to-text), Dedicated Endpoints (reserved GPU capacity on A100, H100, H200, and B200 instances billed per second with a 99.99% uptime SLA), and Friendli Container (self-hosted inference software for customer-managed environments with metering). Layered on top is InferenceSense, a Kubernetes-based platform that lets neocloud operators monetize idle GPU capacity by running FriendliAI inference workloads under a revenue-share model. Distribution combines a self-serve PLG developer motion (570,000+ deployable Hugging Face models with one-click deployment, gated ebook lead capture, and an Anthropic-compatible API), enterprise field sales (custom rate limits, reserved capacity, VPC/on-prem, named CSM), and channel partnerships with AWS, Oracle Cloud Infrastructure, Samsung Cloud Platform, and Nebius.
The business model is primarily usage-based (token-billed serverless inference, per-second GPU billing) supplemented by subscription-based Container licensing and Enterprise contracts. The customer base spans AI startups and AI-native SaaS firms (Upstage, Scatter Lab, NextDay AI, Twelve Labs), large Korean conglomerates and telcos (LG AI Research, SK Telecom), and developer/SMB long-tail users reached through the Hugging Face Hub and self-serve console. The company has raised approximately $26.75M in aggregate disclosed funding (an 8 billion KRW Series A in December 2021 and a $20M seed extension in August 2025), both led by Capstone Partners with Korean institutional co-investors (KB Investment, KDB, KB Securities) and US-based Sierra Ventures and Alumni Ventures. Operations are dual-hub: a 7,000-square-foot commercial and engineering office opened in San Francisco's SoMa district in May 2026 alongside the Seoul engineering hub in Gangnam.
FriendliAI firmographics
Firmographics- Name
- FriendliAI
- Legal name
- FriendliAI Corp.
- Website
- https://friendli.ai
- Company type
- Private
- Founded year
- 2021
- Operating status
- Operating
- Headcount range
- 11–50 employees
- Short description
- FriendliAI is a frontier AI inference cloud that deploys, scales, and monitors large language and multimodal models for AI startups, SaaS firms, and large enterprises. Its optimized GPU serving stack claims up to 3x faster inference and multi-cloud distribution via AWS, OCI, Samsung, and Nebius.
- Ownership category
- akta.pro rank
FriendliAI industry classification
Industry- Product category
- AI Inference Cloud Infrastructure
- NAICS
- Computer Systems Design and Related Services (54151)
- SIC
- Services-Computer Programming, Data Processing, Etc. (7370)
- akta.pro primary industry
- AI Compute Virtualization & Scheduling (GPU virtualization, cluster schedulers) (HDAAAAAG)
Keywords
Where FriendliAI is headquartered
LocationHeadquarters
- HQ city
- Redwood City
- HQ country
- United States
- HQ region
- North America
Offices2 records
Markets served
FriendliAI business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Technology or R&D, Infrastructure, Personnel, Marketing or Sales, Operations
Revenue model
- Model APIs (serverless token-based inference): Pay-per-token consumption pricing across text/vision and speech-to-text models (e.g., GLM-5.2, Llama-3.3-70B, DeepSeek-V3.2, MiniMax-M2.5, K-EXAONE, Whisper-large-v3). Revenue scales with token consumption.
- Dedicated Endpoints (GPU-based inference): On-demand GPU compute billed per second at hourly rates (A100 80GB $2.9/hr, H100 80GB $3.9/hr, H200 141GB $4.5/hr, B200 180GB $8.9/hr). Reserved capacity and On-Demand Endpoints available via the Product Management Console.
- Friendli Container: Customer-deployed container inference serving system priced via sales engagement with per-second metering; revenue is usage-based with subscription-style access terms.
- Enterprise plan: Customizable enterprise framework offering custom Model API rate limits, priority access to high-demand GPU types, reserved GPU capacity, custom region / VPC / on-prem deployments, dedicated support channels, named Customer Success ownership, and custom commercial terms.
- Free Trial / Beta services: Limited-duration free trial periods and beta/eval/pre-release features offered with no charge for evaluation, with auto-conversion to paid metered or subscription billing thereafter.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Usage-based | Pay-as-you-go | Model APIs – pay per 1M tokens (text/vision) or per audio minute (STT) |
| Usage-based | Pay-as-you-go | Dedicated Endpoints – on-demand GPU compute billed per second |
| Subscription | Multi-year contract | Friendli Container – quote-based deployment in customer's environment |
| Subscription | Multi-year contract | Enterprise plan – customizable framework with custom commercial terms |
| Freemium | Pay-as-you-go | Free Trial / Beta services |
Go-to-market motion6 records
Distribution channels6 records
Marketing channels10 records
FriendliAI product offering
Product offeringCore offering
FriendliAI provides a frontier AI inference cloud that lets customers deploy, serve, and scale large language and multimodal models. The platform is delivered through the Friendli Suite console and includes serverless Model APIs billed per token, Dedicated Endpoints billed per second on named GPUs (A100, H100, H200, B200), a self-hosted Friendli Container for in-environment inference, and an InferenceSense add-on that monetizes idle GPU capacity on neocloud infrastructure.
Product overview
FriendliAI operates 'The Frontier AI Inference Cloud,' a platform-plus-modules architecture centered on the Friendli Suite. The Suite is the unified console and orchestration layer through which customers access the three core commercial offerings: Model APIs (serverless, token-priced inference for frontier open-weight and custom models), Dedicated Endpoints (reserved-GPU inference with SLA guarantees and per-second billing), and Friendli Container (self-hosted container inference for customer-controlled environments). Layered on top, InferenceSense is an add-on module that lets neocloud operators monetize idle GPU capacity, while PeriFlow refers to the company's earlier AI development platform. Together, these products let AI engineers deploy frontier open-weight and custom models with continuous batching, speculative decoding, Host KV Cache, and multi-cloud scaling, claiming up to 3x faster inference than vLLM and 50–90% cost savings versus closed model APIs.
Differentiator
Problem solved
Functional benefit
Brands
- Friendli Suite: The umbrella platform name for FriendliAI's inference products including Model APIs, Dedicated Endpoints, Serverless Endpoints, and Container.
- Friendli Dedicated Endpoints
- Friendli Serverless Endpoints
- Friendli Container
- Friendli Model APIs
- PeriFlow
- InferenceSense
Products and services
- Friendli Suite Central management platform orchestrating access to Model APIs, Dedicated Endpoints, and Friendli Container services, with the Product Management Console at suite.friendli.ai, role-based team workspaces, deployment creation, monitoring, metered and subscription billing, and 99.99% uptime targets across geo-distributed infrastructure.
- Model APIs (Friendli Serverless Endpoints) Serverless, pay-per-token inference API providing instant access to frontier open-weight models (text/vision and speech-to-text) without infrastructure setup, with Anthropic Messages API compatibility and one-click deployment of 570,000+ Hugging Face models, billed per million tokens or per audio minute.
- Dedicated Endpoints (Friendli Dedicated Endpoint Service) Reserved-GPU inference service for production-scale workloads, billed per second on A100 80GB, H100 80GB, H200 141GB, and B200 180GB GPUs, with SLA-backed performance guarantees, autoscaling, Host KV Cache, and draft-model speculative decoding for low-latency, high-throughput inference at scale.
- Friendli Container Containerized inference serving system that customers download and run within their own internal environment (cloud or on-prem) with full control over data and deployment, with Metering Server usage tracking and access via customer internal network or VPN.
- InferenceSense Platform enabling neocloud operators to monetize unused GPU capacity by dynamically running inference workloads and sharing revenues with FriendliAI; built on continuous batching and Kubernetes to maximize token throughput on idle GPU fleets.
Companies that use FriendliAI
Customer profileNamed customers6 records
Segments5 records
Ideal customer profiles4 records
FriendliAI technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration6 records
AI capability14 records
Feature10 records
FriendliAI partnerships and signals
Strategic signalPartnerships
Six partnerships are on record, tiered flagship and strategic.
- Samsung Cloud Platform (Samsung SDS)flagshipStrategic alliance to power frontier-model AI inference on NVIDIA B300 GPUs globally. Combines FriendliAI's inference optimization with Samsung's scalable GPU cloud to run frontier open-weight models (GLM-5.1, MiniMax M2.5, NVIDIA Nemotron 3 Super, DeepSeek v3.2) at maximum performance with competitive token pricing. Primarily serves South Korean customers.
- NVIDIAflagshipNVIDIA Nemotron Nano 3 models available day-0 on FriendliAI's inference platform for faster and more cost-efficient AI deployments. Partnership extends to Nemotron 3 Ultra, Nemotron 3 Nano Omni, and host of frontier open-weight models (GLM-5.1, DeepSeek V4, Kimi K2.6) running on NVIDIA GPUs.
- Nebius AI CloudstrategicIntegration that delivers ultra-low latency and cost-efficient AI inference for enterprise customers, supporting over 460,000 models with up to 90% GPU cost savings.
- Hugging FaceflagshipStrategic partnership integrating FriendliAI's inference infrastructure directly into the Hugging Face Hub, enabling developers to deploy generative AI models with a single click on NVIDIA H100 GPUs. Deepens an existing collaboration.
- Amazon Web Services (AWS)flagshipCloud provider partnership that supports FriendliAI inference workloads via AWS, extending distribution to AWS's enterprise customer base globally.
- Oracle Cloud Infrastructure (OCI)flagshipCloud provider partnership that supports FriendliAI inference workloads via OCI, extending distribution to Oracle's enterprise customer base globally.
Scale indicators10 records
Recent moves8 records
Expansion highlights6 records
FriendliAI competitors and assessment
Company assessmentDirect peers
- Together AI: Direct competitor offering a GPU inference cloud for open-weight LLMs with comparable per-token and dedicated deployment pricing, serving AI-native startups and enterprises with similar continuous-batching-style optimizations.
- Fireworks AI: Direct competitor providing serverless and dedicated LLM inference with custom kernels and speculative decoding, targeting AI startups and enterprises with multi-model deployment — overlapping directly with FriendliAI's Model APIs and Dedicated Endpoints.
- Anyscale: Direct peer offering production-scale AI compute and inference on Ray-based infrastructure for LLMs and generative AI workloads, addressing the same AI engineer buyer with comparable performance and scalability claims.
- Modal Labs: Direct peer providing serverless GPU compute and inference for AI workloads, targeting AI startups and developer-led teams with API-first deployment — closely aligned with FriendliAI's self-serve/PLG motion.
- Replicate: Direct competitor running a cloud for open-source generative AI models with per-prediction and dedicated GPU pricing, serving a similar developer-first audience and competing for the same Hugging Face model deployment workload.
- DeepInfra: Direct peer offering low-cost, low-latency inference APIs for open-source LLMs with per-token pricing — comparable model catalog, pricing structure, and target developer persona to FriendliAI Model APIs.
- RunPod: Direct competitor providing GPU cloud and serverless inference endpoints for AI workloads, with comparable per-second GPU billing on H100/A100 SKUs and overlapping AI startup and SMB target customers.
Broad incumbents
- Hugging Face Inference Endpoints: Broader incumbent and direct distribution partner operating its own dedicated inference endpoint service for the same 570,000+ models catalog — competes with FriendliAI while also serving as its primary distribution channel via the Hub partnership.
- AWS Bedrock: Broader incumbent offering managed inference for foundation models inside the AWS ecosystem, competing with FriendliAI for enterprise inference workloads while also serving as a cloud distribution channel.
Emerging players
- OctoAI (acquired by NVIDIA): Emerging player in optimized AI inference infrastructure now operating inside NVIDIA's stack, competing with FriendliAI on price-performance for open-weight model deployment and accelerating the broader inference commoditization narrative.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat6 records
Key risks6 records
Key highlights7 records
Customer concentration
FriendliAI social profiles
Digital presenceFriendliAI compliance and trust
Trust signalCompliance4 records
FriendliAI financial estimates
Financial estimateRevenue estimate
Valuation estimate
FriendliAI leadership team
Management profileNumber of profiles
Profiles3 records
FriendliAI funding detail
Funding detailFunding overview
Funding rounds2 records
Investors6 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
FriendliAI M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about FriendliAI
What does FriendliAI do?
FriendliAI provides a frontier AI inference cloud that lets customers deploy, serve, and scale large language and multimodal models. The platform is delivered through the Friendli Suite console and includes serverless Model APIs billed per token, Dedicated Endpoints billed per second on named GPUs (A100, H100, H200, B200), a self-hosted Friendli Container for in-environment inference, and an InferenceSense add-on that monetizes idle GPU capacity on neocloud infrastructure.
Is FriendliAI a public or private company?
FriendliAI is a private company. It is classified as venture growth investor backed and is currently operating.
When was FriendliAI founded?
FriendliAI was founded in 2021. It employs 11 to 50 people.
Where is FriendliAI based?
FriendliAI is headquartered in Redwood City, United States, in the North America region.
How does FriendliAI make money?
Five revenue lines are on record. Model APIs (serverless token-based inference) is the primary driver. The others are dedicated Endpoints (GPU-based inference), friendli Container, enterprise plan and free Trial / Beta services.
Who are FriendliAI's main competitors?
Direct peers on record are Together AI, Fireworks AI, Anyscale, Modal Labs, Replicate, DeepInfra and RunPod. Broad incumbents are Hugging Face Inference Endpoints and AWS Bedrock. OctoAI (acquired by NVIDIA) is listed as an emerging player.
Does FriendliAI have an API?
Yes. FriendliAI offers public REST/CLI-accessible inference APIs through two primary products: Model APIs (serverless, usage-based inference endpoints for frontier open-weight and custom AI models) and Dedicated Endpoints APIs (reserved GPU capacity with SLA-backed performance targets). The platform also supports compatibility with the Anthropic Messages API specification, enabling developers to integrate FriendliAI inference into existing Anthropic-API-compatible client code. Customer access is governed by API keys/access tokens and configured via the Friendli Suite console (suite.friendli.ai), with service-level commitments specified in the Friendli Suite Service Level Agreement. Developer documentation is at friendli.ai/docs.
What industry is FriendliAI in?
FriendliAI's product category is AI Inference Cloud Infrastructure. Its primary akta.pro industry code is HDAAAAAG, AI Compute Virtualization & Scheduling (GPU virtualization, cluster schedulers). Its NAICS code is 54151 and its SIC code is 7370.