SGLang
SGLang is an open-source LLM and multimodal inference framework from UC Berkeley, hosted under non-profit LMSYS. It powers 400k+ GPUs across NVIDIA, AMD, TPU, Intel, Ascend, and MUSA hardware, with commercial managed hosting via its RadixArk spinout valued at ~$400M.
- Company typePrivate
- Founded2023
- HeadquartersSan Francisco, United States
- Headcount11–50
- GTM typeB2B
- OfferingSoftware
What SGLang does
SGLang is an open-source, high-performance inference framework for large language models and multimodal models, originated at the UC Berkeley lab in 2023 and hosted under the non-profit LMSYS organization. The framework delivers low-latency, high-throughput inference across single-GPU to distributed-cluster deployments through proprietary techniques including RadixAttention (prefix caching), speculative decoding (DFlash, Spec V2), and multi-GPU parallelism. It serves AI/ML developers and enterprise infrastructure teams building production LLM and multimodal applications, supporting Llama, Qwen, DeepSeek, and Nemotron model families with OpenAI API and Hugging Face compatibility, and running natively on NVIDIA, AMD, Intel Xeon, Google TPU, Huawei Ascend NPU, and Moore Threads MUSA hardware.
The product surface extends beyond core text inference. Specialized sub-products include SGLang Diffusion (image generation), SGLang-JAX (TPU/MoE optimization via JAX/Pallas), SGLang-Omni (speech synthesis via Higgs Audio v3 TTS), and Miles, a reinforcement learning framework operated by RadixArk. The commercial arm, RadixArk, spun out from UC Berkeley in January 2026 with Accel-led funding at approximately $400M valuation and offers paid managed hosting for SGLang-based inference and the Miles RL framework, while the open-source engine itself remains free under LMSYS. Distribution is community-first through GitHub, pip, Docker, and Hugging Face, with active engagement across Slack, Discord, X, LinkedIn, weekly public meetings, and a public roadmap.
In production, SGLang is deployed across more than 400,000 GPUs globally, generating trillions of tokens per day, and has been included in NVIDIA CEO Jensen Huang's list of 103 AI-native companies presented at GTC. Named enterprise integrations include Nebius AI Cloud, where SGLang delivered a 2x throughput improvement on DeepSeek R1 inference, alongside joint development with Moreh on an AMD-based distributed inference system and the MUSA backend merge with Moore Threads for native China-market GPU support. Revenue at the open-source level is zero, and monetization flows through RadixArk's commercial operations, which remain in early-stage ramp.
SGLang firmographics
Firmographics- Name
- SGLang
- Legal name
- SGLang (Open-Source Project)
- Website
- https://docs.sglang.ai
- Company type
- Private
- Founded year
- 2023
- Operating status
- Operating
- Headcount range
- 11–50 employees
- Short description
- SGLang is an open-source LLM and multimodal inference framework from UC Berkeley, hosted under non-profit LMSYS. It powers 400k+ GPUs across NVIDIA, AMD, TPU, Intel, Ascend, and MUSA hardware, with commercial managed hosting via its RadixArk spinout valued at ~$400M.
- Ownership category
- akta.pro rank
SGLang industry classification
Industry- Product category
- LLM Inference Framework
- NAICS
- Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (5182), Custom Computer Programming Services (541511)
- SIC
- Services-Prepackaged Software (7372), Services-Computer Programming, Data Processing, Etc. (7370)
- akta.pro primary industry
- Model Hosting, Serving & Inference Platforms (HDAAACAB)
- akta.pro secondary industry
- GraphQL, gRPC & Modern API Protocols (BPAMAOAK)
Keywords
Where SGLang is headquartered
LocationHeadquarters
- HQ city
- San Francisco
- HQ country
- United States
- HQ region
- North America
Offices1 record
Markets served
SGLang business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Technology or R&D, Personnel, Infrastructure, Operations
Revenue model
- Open Source Distribution: SGLang is free open-source software distributed via GitHub, pip, and Docker. No direct revenue from the open-source framework itself.
- RadixArk Commercial Services: Commercial spinout RadixArk offers paid hosting services for SGLang-based inference, plus Miles reinforcement learning framework as a commercial product. Raised approximately $400M valuation in Accel-led funding round.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Freemium | Pay-as-you-go | Open Source - Free tier |
Go-to-market motion2 records
Distribution channels5 records
Marketing channels9 records
SGLang product offering
Product offeringCore offering
SGLang is an open-source, high-performance serving framework for large language models and multimodal models, designed for low-latency, high-throughput inference from a single GPU to large distributed clusters. Its core capabilities include RadixAttention-based prefix caching, multi-GPU parallelism, speculative decoding, and native support across NVIDIA, AMD, Intel Xeon, Google TPU, and Ascend NPU hardware. The framework is OpenAI- and Hugging Face-compatible and is extended by sub-products for diffusion-based image generation, TPU/MoE workloads, and multimodal speech, plus a commercial RL framework (Miles) offered through the RadixArk spinout.
Product overview
SGLang is a unified open-source inference framework (hosted under the non-profit LMSYS organization at UC Berkeley) that serves as both a standalone LLM inference engine and a platform supporting multiple specialized sub-products and modules. The core SGLang framework provides high-performance text and multimodal LLM serving across diverse hardware (NVIDIA, AMD, Intel, TPU, Ascend NPU) with OpenAI and Hugging Face API compatibility. It is extended by RadixAttention (optimization via prefix caching), SGLang Diffusion (image generation), SGLang-JAX (TPU/MoE optimization via JAX/Pallas), SGLang-Omni (speech/voice via Higgs Audio TTS), and — in its commercial RadixArk form — the Miles reinforcement learning framework. Together, these products cover the full stack from inference serving to training and fine-tuning, across text, code, image, and audio modalities.
Differentiator
Problem solved
Functional benefit
Products and services
- SGLang Core Framework Open-source, high-performance serving framework for large language models and multimodal models, providing low-latency, high-throughput inference from a single GPU to large distributed clusters. Built on RadixAttention, prefix caching, and multi-GPU parallelism, with OpenAI- and Hugging Face-compatible APIs and native support for NVIDIA, AMD, Intel Xeon, Google TPU, and Ascend NPU hardware. Used by AI developers, researchers, and enterprise AI infrastructure teams to deploy Llama, Qwen, DeepSeek, and other major model families.
- SGLang Diffusion Sub-product of the SGLang ecosystem that provides serving infrastructure for diffusion-based image generation models, enabling multimodal workloads beyond text. Targeted at developers and AI infrastructure teams deploying image generation models within the SGLang runtime.
- SGLang-JAX Variant of SGLang optimized for Google TPU hardware via the JAX/Pallas framework. Demonstrated with the Ling-2.6-1T MoE model using a single Pallas kernel that hides MoE data movement behind compute, enabling efficient inference and training of mixture-of-experts models on TPU clusters. Targeted at AI teams running MoE workloads on TPU infrastructure.
- SGLang-Omni Multimodal extension of SGLang that supports speech and voice models, including Higgs Audio v3 TTS for real-time, controllable text-to-speech generation for voice agents. Targeted at developers building voice and conversational AI applications on top of the SGLang runtime.
- Miles Reinforcement Learning Framework Commercial reinforcement learning framework developed by RadixArk (the commercial spinout of the SGLang UC Berkeley team) to complement the open-source SGLang inference engine with RL training and fine-tuning capabilities. Sold alongside RadixArk's paid hosting services to enterprise customers that need to train and fine-tune AI models in the same stack where they run inference.
Quantifiable outcome
- 2x throughput improvement for DeepSeek R1
- +4 more outcomes
Companies that use SGLang
Customer profileNamed customers2 records
Segments2 records
Ideal customer profiles3 records
SGLang technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration2 records
AI capability9 records
Feature6 records
SGLang partnerships and signals
Strategic signalPartnerships
Three partnerships are on record, tiered core.
- Moore ThreadscoreMoore Threads and SGLang co-hosted the SGLang x MUSA Meetup in Beijing on May 10, 2026. This collaboration marks the official merge of the MUSA backend into SGLang mainline, enabling developers to use SGLang directly with Moore Threads GPUs without third-party adaptation layers. Signals China's GPU open-source ecosystem entering native support phase.
- MorehcoreMoreh and SGLang announced joint development of AMD-based distributed inference system at AI Infra Summit 2025. Moreh unveiled distributed inference system on AMD hardware showcasing collaboration with SGLang, expanding into deep learning inference market.
- NebiuscoreNebius AI Cloud collaborated with SGLang to enhance DeepSeek R1 performance, achieving 2x throughput improvement. Nebius highlighted SGLang partnership as key update in May 2025 monthly digest, demonstrating successful integration with cloud AI infrastructure.
Scale indicators4 records
Recent moves6 records
Expansion highlights8 records
SGLang competitors and assessment
Company assessmentDirect peers
- vLLM: vLLM is the closest direct competitor — an open-source high-throughput LLM serving engine with PagedAttention and continuous batching, originating from UC Berkeley-affiliated researchers and competing head-to-head with SGLang for the same open-source developer audience and RadixArk-style commercial paths.
- Hugging Face Text Generation Inference (TGI): Hugging Face's open-source Rust/Python LLM serving framework — a direct peer that targets the same open-source deployment audience, integrates tightly with the Hugging Face model hub, and competes for community mindshare in self-hosted LLM serving.
Emerging players
- Fireworks AI: Fireworks AI is an inference-focused AI cloud provider offering production LLM serving with proprietary optimization techniques — comparable to RadixArk as a managed inference offering while competing for the same enterprise customers that SGLang's commercial spinout targets.
- Replicate: Replicate operates a managed cloud for running open-source ML models — comparable as a hosted-inference business model parallel to RadixArk, though more oriented to a long-tail of models rather than frontier LLM serving.
- Together AI: Together AI operates an open-source-friendly inference cloud (Together Inference) built around its own optimizations and forks of popular engines — an emerging player that combines managed hosting (analogous to RadixArk) with proprietary inference techniques competing with SGLang's commercial offering.
- Anyscale: Anyscale, the commercial company behind Ray and Ray Serve, offers distributed compute and model serving infrastructure — comparable to SGLang's distributed inference capability (multi-GPU parallelism, cluster-scale deployment) with a more general-purpose positioning.
- Modular: Modular builds the Mojo language and MAX inference platform targeting high-performance AI deployment — an emerging player competing in the same inference-optimization space as SGLang with a focus on cross-hardware portability and developer ergonomics.
Others
- LMSYS Org / Chatbot Arena: LMSYS Org is the non-profit parent that hosts SGLang and operates Chatbot Arena — thematically related as the governance vehicle and ecosystem brand under which SGLang is developed, but not a commercial competitor.
- DeepSeek: DeepSeek is a frontier AI lab whose open-weight models (e.g., DeepSeek R1, V3) are among the most heavily served by SGLang (per Nebius case study). Not a direct competitor but a critical ecosystem dependency — model popularity drives SGLang adoption.
Broad incumbents
- NVIDIA TensorRT-LLM: NVIDIA's first-party LLM inference optimization library — a broad incumbent that ships natively with NVIDIA hardware and competes with SGLang for high-performance GPU inference workloads, particularly where deep NVIDIA tooling integration matters.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat5 records
Key risks6 records
Key highlights7 records
Customer concentration
SGLang social profiles
Digital presenceSGLang financial estimates
Financial estimateRevenue estimate
Valuation estimate
SGLang leadership team
Management profileNumber of profiles
SGLang funding detail
Funding detailFunding overview
Funding rounds
Investors
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
SGLang M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about SGLang
What does SGLang do?
SGLang is an open-source, high-performance serving framework for large language models and multimodal models, designed for low-latency, high-throughput inference from a single GPU to large distributed clusters. Its core capabilities include RadixAttention-based prefix caching, multi-GPU parallelism, speculative decoding, and native support across NVIDIA, AMD, Intel Xeon, Google TPU, and Ascend NPU hardware. The framework is OpenAI- and Hugging Face-compatible and is extended by sub-products for diffusion-based image generation, TPU/MoE workloads, and multimodal speech, plus a commercial RL framework (Miles) offered through the RadixArk spinout.
Is SGLang a public or private company?
SGLang is a private company. It is classified as nonprofit foundation owned and is currently operating.
When was SGLang founded?
SGLang was founded in 2023. It employs 11 to 50 people.
Where is SGLang based?
SGLang is headquartered in San Francisco, United States, in the North America region.
How does SGLang make money?
Two revenue lines are on record. Open Source Distribution is the primary driver. The others are radixArk Commercial Services.
Who are SGLang's main competitors?
Direct peers on record are vLLM and Hugging Face Text Generation Inference (TGI). Emerging players are Fireworks AI, Replicate, Together AI, Anyscale and Modular. Others are LMSYS Org / Chatbot Arena and DeepSeek. NVIDIA TensorRT-LLM is listed as a broad incumbent.
Does SGLang have an API?
Yes. SGLang provides OpenAI API-compatible endpoints, including /v1/chat/completions and /v1/embeddings, enabling drop-in compatibility for applications built against the OpenAI API. It also supports Hugging Face model formats, allowing model deployment with minimal code changes. JSON endpoints for /v1/rerank are also exposed. The API is served via a model server launched via the SGLang runtime. Developer documentation is at docs.sglang.io.
What industry is SGLang in?
SGLang's product category is LLM Inference Framework. Its primary akta.pro industry code is HDAAACAB, Model Hosting, Serving & Inference Platforms, with a secondary code of BPAMAOAK, GraphQL, gRPC & Modern API Protocols. Its NAICS code is 5182 and its SIC code is 7372.