Novita AI
Novita AI is a San Francisco-based AI-native cloud platform that provides serverless model APIs for 200+ LLMs and generative media models, Firecracker-microVM-based Agent Sandbox runtime for autonomous AI agents, and GPU cloud infrastructure (H100, H200, B200, RTX 5090/4090), targeting developers and AI builders via a self-serve platform and a strategic Hugging Face distribution partnership.
- Company typePrivate
- Founded2024
- HeadquartersSan Francisco, United States
- Headcount11–50
- GTM typeB2B
- OfferingSoftware
What Novita AI does
Novita AI (San Francisco, founded 2024) operates an AI-native cloud platform that unifies three product pillars: serverless Model APIs exposing 200+ large language, image, video, audio, vision, and code models through OpenAI- and Anthropic-compatible endpoints; Agent Sandbox, a Firecracker microVM-based runtime providing kernel-isolated execution environments for autonomous coding and browser-use agents with sub-200ms startup and stateful pause/resume; and a GPU Cloud spanning dedicated GPU instances (H100, H200, B200, RTX 5090/4090), serverless GPU jobs that scale to zero, and bare-metal clusters with NVLink and 400 Gb/s RDMA interconnect. Dedicated Endpoints (Standard 98% SLA / Pro 99.5% SLA) sit alongside for enterprise production workloads.
The business model is predominantly usage-based — per-token API billing, hourly GPU rentals, per-second sandbox execution, and per-asset media generation — with a subscription layer for Dedicated Endpoints and an affiliate channel paying 10% commissions for 180 days. Pricing is positioned at up to 50% below major cloud providers (H100 SXM at $1.70/GPU/hr, B200 SXM at $4.77/GPU/hr). Go-to-market is product-led growth on a self-serve platform (sign-up includes $100 sandbox credits, no credit card required), reinforced by a Discord developer community, 50+ framework integrations (Hugging Face, LlamaIndex, LangChain, vLLM, SGLang, Claude Code, Dify, Continue, etc.), a strategic partnership with Hugging Face exposing Novita to 5M+ developers via a 'Deploy on Novita' flow, and direct enterprise sales for higher-SLA tiers. The company holds AICPA SOC 2 Type 2 certification and reports serving 350,000+ developers on the platform.
Customer concentration is mixed across named logos (Hugging Face, Quora/POE, OpenRouter, Vercel, Fish Audio, Kilo Code, Genspark, Gizmo, TiDB, Hygo, Wiz) spanning developer-tools, AI infrastructure, search, and audio verticals. The model API business is essentially a commoditized inference reseller exposed to GPU cost and competitive price compression, while the Agent Sandbox line and the Hugging Face channel partnership are the more differentiated, defensible assets. Headcount of 1–10 employees against the operational footprint claimed (1,000+ GPU H100 cluster, 24/7 inference, SOC 2 controls) implies either heavy reliance on automated/managed operations or significant under-reporting of true workforce.
Novita AI firmographics
Firmographics- Name
- Novita AI
- Legal name
- Novita AI
- Website
- https://novita.ai
- Company type
- Private
- Founded year
- 2024
- Operating status
- Operating
- Headcount range
- 11–50 employees
- Short description
- Novita AI is a San Francisco-based AI-native cloud platform that provides serverless model APIs for 200+ LLMs and generative media models, Firecracker-microVM-based Agent Sandbox runtime for autonomous AI agents, and GPU cloud infrastructure (H100, H200, B200, RTX 5090/4090), targeting developers and AI builders via a self-serve platform and a strategic Hugging Face distribution partnership.
- Ownership category
- akta.pro rank
Novita AI industry classification
Industry- Product category
- AI Cloud Infrastructure
- NAICS
- Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (5182), Software Publishers (513210), Computer Systems Design and Related Services (54151)
- SIC
- Services-Prepackaged Software (7372), Services-Computer Integrated Systems Design (7373), Services-Computer Processing & Data Preparation (7374)
- akta.pro primary industry
- Model Hosting, Serving & Inference Platforms (HDAAACAB)
- akta.pro secondary industries
- AI Compute Virtualization & Scheduling (GPU virtualization, cluster schedulers) (HDAAAAAG), Responsible AI, Security & Privacy Platforms (Safety, Guardrails, PII) (HDAEANAG)
Keywords
Where Novita AI is headquartered
LocationHeadquarters
- HQ city
- San Francisco
- HQ country
- United States
- HQ region
- North America
Offices1 record
Markets served
Novita AI business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Technology or R&D, Infrastructure, Personnel, Marketing or Sales, Operations
Revenue model
- Serverless Model API (Pay-per-token): Customers pay per token consumed (input and output tokens) for LLM, image, audio, and video generation via serverless endpoints. No infrastructure to manage. Supports batch inference at 50% discount on input/output tokens. Revenue scales with customer usage.
- GPU Instance Rentals: Dedicated GPU machines (H100 SXM at $1.70/GPU/hr, B200 SXM at $4.77/GPU/hr, RTX 5090, RTX 4090) with hourly billing. Customers fully control the infrastructure with predictable performance and no shared resources.
- GPU Bare Metal: Reserved physical GPU clusters for large-scale inference, training, and enterprise deployments with contractual SLAs and guaranteed delivery. Billed hourly at premium rates.
- Dedicated Endpoints: Private model endpoints with guaranteed performance and isolated resources. Priced on a monthly subscription basis with Standard (98% SLA) and Pro (99.5% SLA) tiers.
- Agent Sandbox: Secure isolated runtimes for AI agents. Billing per second for sandbox execution time (vCPU hours, RAM hours, storage). Available via managed NovitaClaw deployments, Sandbox Skills for OpenClaw, and Hermes-compatible runtimes.
- Affiliate Program: Partners earn 10% commission on referral spending for the first 180 days. Commission is a cost acquisition channel for Novita.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Usage-based | Pay-as-you-go | LLM Serverless API - usage-based per token |
| Usage-based | Pay-as-you-go | GPU Instance rentals - hourly rate |
| Usage-based | Pay-as-you-go | GPU Bare Metal - hourly rate |
| Usage-based | Pay-as-you-go | Agent Sandbox - billed per second |
| Subscription | Monthly | Dedicated Endpoints - monthly subscription |
| Unit Pricing | Pay-as-you-go | Image Generation - per image |
| Unit Pricing | Pay-as-you-go | Video Generation - per video or per second |
| Unit Pricing | Pay-as-you-go | Audio/TTS - per million characters or per voice |
Go-to-market motion3 records
Distribution channels5 records
Marketing channels8 records
Novita AI product offering
Product offeringCore offering
Novita AI is an AI-native cloud platform that runs 200+ AI models (LLMs, image, audio, video, vision) through OpenAI and Anthropic-compatible APIs, scales dedicated and serverless GPU compute (H100, H200, B200, RTX 5090/4090), and provides a Firecracker microVM-based Agent Sandbox for securely executing autonomous AI agents. Revenue is generated via per-token API billing, hourly GPU rental, per-second sandbox execution, and monthly dedicated endpoint subscriptions.
Product overview
Novita AI is an AI-native cloud platform providing a unified infrastructure for running AI models, scaling GPU compute, and building autonomous agents. The platform consists of three core product pillars: (1) Model APIs offering serverless access to 200+ LLMs, image, video, and audio generation models; (2) Agent Sandbox providing secure isolated runtime environments using Firecracker microVMs for autonomous coding agents; and (3) GPU Cloud encompassing GPU Instances, Serverless GPU, and Bare Metal options for flexible infrastructure deployment. Additional offerings include Dedicated Endpoints for private model access with SLA guarantees, NovitaClaw agent framework, and Sandbox Skills modules. The platform is OpenAI and Anthropic API-compatible, enabling day-zero deployment of new model releases with up to 50ms time-to-first-token performance.
Differentiator
Problem solved
Functional benefit
Brands
- NovitaClaw: Managed deployment product for AI agents
- Novita Sandbox
Products and services
- Model APIs Serverless API platform providing access to 200+ AI models spanning LLMs (Deepseek V4 Pro, Qwen3.5, GLM-5.1, Kimi K2.6, Gemma 4, Llama 4, Mistral), image generation (Flux, SDXL, Hunyuan Image 3, Seedream, Qwen Image), video generation (Kling, Hunyuan Video, Wan, MiniMax, PixVerse), audio (Fish Audio TTS, MiniMax speech), and vision models through OpenAI and Anthropic-compatible endpoints, billed per token with no infrastructure management required. Targets developers building AI-powered applications.
- Agent Sandbox Secure runtime infrastructure using Firecracker microVMs that isolates autonomous AI agent systems during execution with individual kernels and isolated memory boundaries, supporting sub-200ms startup times, scaling to thousands of parallel microVMs, and stateful pause/resume for long-running workflows. Designed for coding agents, API-calling agents, and web-browsing agents; available through managed NovitaClaw deployments, Sandbox Skills for OpenClaw, and Hermes-compatible runtimes.
- GPU Instance Full-control dedicated GPU machines available in seconds with NVIDIA H100, H200, A100, L4, RTX 4090, and RTX 5090 configurations. Offers predictable performance with isolated resources, no shared infrastructure, and supports deploying models, running inference, and training from scratch. Targets AI/ML teams needing dedicated GPU compute with hourly billing.
- Serverless GPU On-demand GPU job execution with automatic resource allocation, auto-scaling to zero when idle, and pay-only-for-execution billing model. No instances to provision or idle compute to manage, making it suitable for intermittent or bursty inference and batch workloads.
- GPU Bare Metal Dedicated physical GPU clusters (H100 SXM, B200 SXM, H200 SXM, RTX 5090, RTX 4090) for large-scale inference, training runs, and enterprise deployments requiring maximum performance with zero virtualization overhead. Billed hourly at premium rates (H100 SXM at $1.70/GPU/hr, B200 SXM at $4.77/GPU/hr) with 1,000+ GPU linear scaling on H100 SXM clusters.
- Dedicated Endpoints Private model endpoints with guaranteed performance and isolated resources ensuring consistent latency at any throughput without noisy-neighbor issues. Available in Standard (98% SLA) and Pro (99.5% SLA) tiers on a monthly subscription basis, designed for production AI deployments requiring SLA-backed reliability.
- GPU Cloud Unified GPU infrastructure platform encompassing GPU Instances, Serverless GPU, and Bare Metal offerings, providing flexible deployment options from fully managed serverless to dedicated physical hardware. Includes multi-node GPU clusters with NVLink 4th Gen, GPUDirect RDMA, and 400 Gb/s RDMA networking (e.g., CLUSTER-01 with 6x H200 nodes, CLUSTER-02 with 6x H100 nodes) for different workload requirements.
Quantifiable outcome
- Up to 50% cost savings compared to major cloud providers
- +6 more outcomes
Companies that use Novita AI
Customer profileNamed customers9 records
Segments4 records
Ideal customer profiles5 records
Novita AI technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration16 records
AI capability15 records
Feature5 records
Novita AI partnerships and signals
Strategic signalPartnerships
Seven partnerships are on record, tiered core and minor.
- Hugging FacecoreNovita AI and Hugging Face announced a strategic partnership enabling over five million developers on the Hugging Face platform to instantly deploy AI models as production-ready APIs without infrastructure setup. The partnership introduces a seamless 'Deploy on Novita' experience removing operational overhead. Novita AI was a day 0 launch partner for Google's Gemma 4 model, and is also available directly on Hugging Face's platform.
- POEminorNovita models are available on POE (Quora's AI platform), enabling POE's user base to access Novita's inference infrastructure.
- vLLMcoreNovita AI partnered with vLLM to advance AI inference, optimizing the vLLM inference engine on Novita's GPU infrastructure for better performance and cost efficiency.
- SGLangcoreIntegration with SGLang (Structured Generation Language) framework, enabling developers to build and deploy AI applications using SGLang on Novita's GPU cloud.
- LlamaIndexcoreOfficial Novita AI integration with LlamaIndex, enabling RAG (Retrieval Augmented Generation) workflows using Novita's model APIs.
- LangchaincoreOfficial integration with LangChain framework, allowing LangChain developers to easily connect to Novita's model APIs and GPU infrastructure.
- Claude (Anthropic)coreClaude Code integration guide published by Novita, enabling developers to use Claude Code with Novita's infrastructure for AI-powered coding workflows.
Scale indicators12 records
Recent moves6 records
Expansion highlights6 records
Novita AI competitors and assessment
Company assessmentBroad incumbents
- CoreWeave: Large-scale GPU cloud provider originally built for crypto/HPC workloads, now serving major AI labs with H100/B200 clusters and bare metal. Competes with Novita's Bare Metal and GPU Cluster products but operates at substantially larger scale with enterprise-only GTM.
- Lambda Labs: GPU cloud and on-prem AI infrastructure provider offering H100 clusters, GPU cloud instances, and model API services. Overlaps with Novita's GPU Cloud and Bare Metal offerings but with a heavier on-prem and training orientation.
- Crusoe: Vertically-integrated GPU cloud built around stranded energy, serving AI training and inference workloads with H100 clusters. Comparable to Novita's Bare Metal and GPU Cluster offerings; co-listed with Novita on Ramp's trending AI infrastructure list, signaling similar customer overlap.
Direct peers
- Replicate: Cloud API for running open-source ML models (LLMs, image, video, audio) on demand with per-second billing. Directly comparable to Novita's serverless multi-modal model API and similar developer-led GTM motion.
- RunPod: GPU cloud offering on-demand instances and serverless endpoints for AI training and inference at consumer-friendly hourly rates. Directly comparable to Novita's GPU Instance and Serverless GPU products for individual developers and small teams.
- Anyscale: Commercial platform behind the open-source Ray distributed computing framework, offering managed compute for AI/ML workloads including LLM serving. Comparable to Novita's GPU Cloud and vLLM-optimized inference stack for production AI deployments.
- Modal Labs: Serverless cloud platform for running AI/ML workloads with Python-based infrastructure-as-code and per-second GPU billing. Overlaps with Novita's Serverless GPU and Agent Sandbox offerings for developers running code-driven AI pipelines.
- Fireworks AI: Production inference platform for open and proprietary LLMs with serverless APIs and dedicated deployments, focused on low-latency, cost-efficient serving. Competes head-to-head with Novita on the same token-priced model API product.
- Together AI: Open-source-focused AI inference cloud offering serverless model APIs and dedicated GPU instances across 200+ open models. Closest direct competitor to Novita's model API and GPU cloud pillars, with comparable OpenAI-compatible endpoints and similar price-performance positioning.
Others
- Hugging Face: While Novita's largest distribution partner and customer (not direct competitor), Hugging Face Inference Endpoints offers competing dedicated inference infrastructure. Critical ecosystem anchor for Novita's developer acquisition and a potential competitor in the inference layer.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat6 records
Key risks6 records
Key highlights7 records
Customer concentration
Novita AI social profiles
Digital presenceNovita AI compliance and trust
Trust signalCompliance1 record
Novita AI financial estimates
Financial estimateRevenue estimate
Valuation estimate
Novita AI leadership team
Management profileNumber of profiles
Novita AI funding detail
Funding detailFunding overview
Funding rounds
Investors
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Novita AI M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Novita AI
What does Novita AI do?
Novita AI is an AI-native cloud platform that runs 200+ AI models (LLMs, image, audio, video, vision) through OpenAI and Anthropic-compatible APIs, scales dedicated and serverless GPU compute (H100, H200, B200, RTX 5090/4090), and provides a Firecracker microVM-based Agent Sandbox for securely executing autonomous AI agents. Revenue is generated via per-token API billing, hourly GPU rental, per-second sandbox execution, and monthly dedicated endpoint subscriptions.
Is Novita AI a public or private company?
Novita AI is a private company. It is classified as unknown and is currently operating.
When was Novita AI founded?
Novita AI was founded in 2024. It employs 11 to 50 people.
Where is Novita AI based?
Novita AI is headquartered in San Francisco, United States, in the North America region.
How does Novita AI make money?
Six revenue lines are on record. Serverless Model API (Pay-per-token) is the primary driver. The others are GPU Instance Rentals, GPU Bare Metal, dedicated Endpoints, agent Sandbox and affiliate Program.
Who are Novita AI's main competitors?
Broad incumbents on record are CoreWeave, Lambda Labs and Crusoe. Direct peers are Replicate, RunPod, Anyscale, Modal Labs, Fireworks AI and Together AI. Hugging Face is listed as an others.
Does Novita AI have an API?
Yes. Novita AI provides a comprehensive API platform offering OpenAI-compatible and Anthropic-compatible endpoints for LLM inference, image generation, video generation, audio processing, and model fine-tuning. The platform supports serverless and dedicated endpoint deployments with authentication via Bearer token. API documentation is available at https://novita.ai/docs with endpoints documented for GPU instances, model APIs, and sandbox management. MCP (Model Context Protocol) server is available as indicated by the Agent Discovery Metadata listing /.well-known/mcp/server-card.json. Developer documentation is at novita.ai/docs.
What industry is Novita AI in?
Novita AI's product category is AI Cloud Infrastructure. Its primary akta.pro industry code is HDAAACAB, Model Hosting, Serving & Inference Platforms, with a secondary code of HDAAAAAG, AI Compute Virtualization & Scheduling (GPU virtualization, cluster schedulers). Its NAICS code is 5182 and its SIC code is 7372.