BentoML
BentoML provides a unified AI inference platform for packaging, deploying, and scaling machine learning and LLM models across any cloud or on-premises environment. It serves AI-first enterprises and developers across digital presence, navigation, financial services, travel, and e-commerce.
- Company typePrivate
- Founded2019
- HeadquartersSan Francisco, United States
- Headcount11–50
- GTM typeB2B
- OfferingSoftware
What BentoML does
BentoML (legal entity Atalaya Tech, Inc.) is an AI inference infrastructure company headquartered in San Francisco. It provides a unified platform for packaging, deploying, scaling, and observing machine learning and large language model (LLM) inference workloads. The core offering, the Bento Inference Platform, includes the Bento Compute Engine for elastic auto-scaling, scale-to-zero, cold-start acceleration, and multi-cloud orchestration; an Open Model Catalog for deploying frontier open-source models such as Llama 4, DeepSeek, GPT-OSS, and Qwen with pre-configured optimization; Custom Model Serving supporting vLLM, TRT-LLM, JAX, SGLang, PyTorch, and Transformers; an LLM Gateway that abstracts multiple model providers; and Bring Your Own Cloud (BYOC) deployment on AWS, GCP, or Azure for security- and privacy-sensitive customers. The open-source BentoML framework on GitHub anchors a product-led growth motion and feeds the managed BentoCloud service.
The company serves AI-first enterprises and developers deploying multiple models in production across verticals including digital presence (Yext), navigation and maps (TomTom, Naver, Koo), financial services (Mission Lane, AdThena), travel (GetYourGuide), and e-commerce (Shopback). Reported customer outcomes include up to 90% compute cost reduction (fintech), ~70% development time savings, 2-3x faster deployment cycles, and scale-to-zero economics for early-stage teams like Neurolabs. Revenue is generated through usage-based GPU compute on BentoCloud, enterprise subscriptions with forward-deployed engineering and SLAs, and indirectly from open-source adoption that funnels into cloud conversions. Distribution combines self-serve (GitHub, cloud.bentoml.com, AWS/GCP marketplaces) with field sales for enterprise.
In 2024, BentoML announced it was joining Modular, a strategic operating company focused on AI infrastructure, to combine its inference platform with Modular's AI compiler technology and build the next generation of AI inference infrastructure. Pricing for the combined offering now routes through Modular's pricing structure. BentoML has raised approximately $19M across three priced rounds (2020, 2023, 2025) led by Bow Capital, DCM Ventures, and Alpha Intelligence Capital.
BentoML firmographics
Firmographics- Name
- BentoML
- Legal name
- Atalaya Tech, Inc.
- Website
- https://bentoml.com
- Company type
- Private
- Founded year
- 2019
- Operating status
- Operating
- Headcount range
- 11–50 employees
- Short description
- BentoML provides a unified AI inference platform for packaging, deploying, and scaling machine learning and LLM models across any cloud or on-premises environment. It serves AI-first enterprises and developers across digital presence, navigation, financial services, travel, and e-commerce.
- Ownership category
- akta.pro rank
BentoML industry classification
Industry- Product category
- AI Inference Platform
- NAICS
- Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (5182)
- SIC
- Services-Prepackaged Software (7372)
- akta.pro primary industry
- Model Deployment, Serving & Inference Platforms (HDAAABAF)
- akta.pro secondary industries
- Model Hosting, Serving & Inference Platforms (HDAAACAB), Model Compression & Optimization (Quantization, Distillation, Pruning) (HDAAACAL), Model Serving, Inference & Deployment Platforms (APIs, Edge/On-Prem) (HDAEANAC)
Keywords
Where BentoML is headquartered
LocationHeadquarters
- HQ city
- San Francisco
- HQ country
- United States
- HQ region
- North America
Offices1 record
Markets served
BentoML business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Technology or R&D, Infrastructure, Personnel, Marketing or Sales, Operations
Revenue model
- Managed Cloud Services: BentoCloud provides managed inference infrastructure with GPU access, auto-scaling, and operational tooling. Revenue generated from compute usage and platform fees for hosted deployment.
- Enterprise Subscriptions: Enterprise-tier subscriptions with dedicated support, forward-deployed engineering, custom optimization, and SLA guarantees for mission-critical deployments.
- Open-Source Framework: Open-source BentoML framework available freely on GitHub. Revenue generated indirectly through community adoption leading to cloud conversions and enterprise deals.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Freemium | Monthly | BentoML Open-Source - Free tier for developers |
| Usage-based | Pay-as-you-go | Bento Inference Platform - Managed cloud services |
| Subscription | Multi-year contract | Enterprise - Dedicated support and custom deployments |
Go-to-market motion2 records
Distribution channels4 records
Marketing channels8 records
BentoML product offering
Product offeringCore offering
BentoML provides a unified AI inference platform for packaging, deploying, and scaling AI/ML models in production. Its offerings include an open-source framework (BentoML on GitHub) for serving any model architecture across frameworks such as vLLM, TRT-LLM, JAX, SGLang, PyTorch, and Transformers, plus a managed cloud platform (Bento Inference Platform / BentoCloud) featuring auto-scaling, scale-to-zero, cold-start acceleration, observability, and Bring-Your-Own-Cloud (BYOC) deployment on AWS, GCP, and Azure.
Product overview
BentoML is an AI inference infrastructure company that helps organizations build and deploy AI applications at scale. Following BentoML's joining of Modular in April 2026, the company now operates as part of Modular's platform, offering tools for deploying various AI models including DeepSeek variants on any cloud infrastructure with features like advanced autoscaling, multi-GPU support, and built-in LLM-specific observability.
Differentiator
Problem solved
Functional benefit
Brands
- Bento Inference Platform: Full control inference platform for self-hosting AI models, deploying anywhere, serving any model with tailored optimization.
- BentoML Open-Source
- LLM Inference Handbook
- LLM Performance Explorer
Products and services
- Bento Inference Platform
Quantifiable outcome
- 70% development time reduction (Yext)
- +7 more outcomes
Companies that use BentoML
Customer profileNamed customers9 records
Segments3 records
Ideal customer profiles3 records
BentoML technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- No
- API docs
- API detail
Core technology
AI maturity
App detail
Feature8 records
BentoML partnerships and signals
Strategic signalPartnerships
Ten partnerships are on record, tiered core, supporting and flagship.
- Hugging FacecoreHugging Face integration for model deployment. BentoML provides optimized deployment path for Hugging Face models with seamless model loading and inference serving capabilities.
- AWScoreAWS Marketplace listing for BentoML deployment. Support for AWS GPU instances (A100, H100) and integration with AWS infrastructure for BYOC deployments.
- Google CloudcoreGoogle Cloud Marketplace listing for BentoML. Support for GCP GPU instances and integration with Google Cloud infrastructure.
- NVIDIAcoreNVIDIA GPU optimization and support across H100, H200, B200, A100, and other NVIDIA data center GPUs. BentoML provides optimized inference configurations for NVIDIA hardware.
- AMDsupportingAMD GPU support including MI300X, MI350X, MI355X for customers preferring AMD hardware. BentoML provides optimization guidance for AMD GPU deployment.
- IntelsupportingIntel technology partnership for AI infrastructure deployment. Joint content and technical collaboration on optimized deployment patterns.
- TwiliosupportingTechnology partnership for voice and communication AI applications. Joint solution for building voice applications with BentoML inference.
- vLLMcoreDeep integration with vLLM open-source inference engine. BentoML provides seamless deployment of vLLM-powered models with optimized configurations.
- SGLangcoreIntegration with SGLang framework for structured LLM deployment. BentoML supports SGLang as an inference backend option.
- ModularflagshipBentoML announced joining Modular in 2025. Partnership combines BentoML's inference platform with Modular's AI compiler technology to build next generation of AI inference infrastructure.
Scale indicators12 records
Recent moves7 records
Expansion highlights6 records
BentoML competitors and assessment
Company assessmentDirect peers
- Anyscale: Anyscale offers a Ray-based AI compute platform with managed inference and serving capabilities for any model. It directly competes with BentoML on managed inference, multi-framework support, and enterprise deployment.
- Fireworks AI: Fireworks AI provides a managed inference platform optimized for open-source and custom LLMs with cost and latency tuning. It competes head-on with BentoML on LLM serving economics and developer-led GTM.
- Together AI: Together AI operates an open-source-focused AI cloud for training, fine-tuning, and serving LLMs. It overlaps directly with BentoML on inference serving of open-source models with usage-based pricing.
- Baseten: Baseten provides a model-serving platform with autoscaling, observability, and deployment of custom models. It competes with BentoML on the developer-to-enterprise inference platform motion.
- Replicate: Replicate offers a cloud platform to run machine learning models via API with usage-based pricing. It serves a similar developer/enterprise audience deploying open-source and custom models in production.
- Modal: Modal provides serverless cloud compute infrastructure purpose-built for AI/ML workloads including inference serving. It competes with BentoCloud on managed GPU compute and developer experience for AI deployment.
- DeepInfra: DeepInfra offers low-cost, low-latency inference-as-a-service for open-source LLMs and other models. It directly competes on managed inference pricing and model coverage.
Broad incumbents
- AWS SageMaker: AWS SageMaker is the hyperscaler incumbent for end-to-end ML training and inference on AWS. It overlaps with BentoML on model deployment and serving, with the advantage of native AWS integration and enterprise procurement.
- Google Vertex AI: Vertex AI is Google Cloud's managed ML platform for training, deploying, and serving models. It competes broadly with BentoML on enterprise model serving within the GCP ecosystem and through multi-cloud portability claims.
- Databricks Mosaic AI: Databricks Mosaic AI provides model serving and MLOps within the Databricks lakehouse platform. It overlaps with BentoML on enterprise inference, particularly for organizations already standardized on Databricks.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat5 records
Key risks5 records
Key highlights6 records
Customer concentration
BentoML social profiles
Digital presenceBentoML compliance and trust
Trust signalCompliance2 records
BentoML financial estimates
Financial estimateRevenue estimate
Valuation estimate
BentoML leadership team
Management profileNumber of profiles
Profiles2 records
BentoML funding detail
Funding detailFunding overview
Funding rounds4 records
Investors9 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
BentoML M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about BentoML
What does BentoML do?
BentoML provides a unified AI inference platform for packaging, deploying, and scaling AI/ML models in production. Its offerings include an open-source framework (BentoML on GitHub) for serving any model architecture across frameworks such as vLLM, TRT-LLM, JAX, SGLang, PyTorch, and Transformers, plus a managed cloud platform (Bento Inference Platform / BentoCloud) featuring auto-scaling, scale-to-zero, cold-start acceleration, observability, and Bring-Your-Own-Cloud (BYOC) deployment on AWS, GCP, and Azure.
Is BentoML a public or private company?
BentoML is a private company. It is classified as corporate owned and is currently operating.
When was BentoML founded?
BentoML was founded in 2019. It employs 11 to 50 people.
Where is BentoML based?
BentoML is headquartered in San Francisco, United States, in the North America region.
How does BentoML make money?
Three revenue lines are on record. Managed Cloud Services are the primary driver. The others are enterprise Subscriptions and open-Source Framework.
Who are BentoML's main competitors?
Direct peers on record are Anyscale, Fireworks AI, Together AI, Baseten, Replicate, Modal and DeepInfra. Broad incumbents are AWS SageMaker, Google Vertex AI and Databricks Mosaic AI.
Does BentoML have an API?
No public API is recorded for BentoML.
What industry is BentoML in?
BentoML's product category is AI Inference Platform. Its primary akta.pro industry code is HDAAABAF, Model Deployment, Serving & Inference Platforms, with a secondary code of HDAAACAB, Model Hosting, Serving & Inference Platforms. Its NAICS code is 5182 and its SIC code is 7372.