Developer docs
API playgroundTry for free, no card

Search company profiles

Novita AI

Full company profile

uuid00059jv

Namestring
Novita AI
Legal namestring
Novita AI
Websiteurl
novita.ai
Company typeenum
Private
Founded yearint
2024
Descriptiontext

Novita AI (San Francisco, founded 2024) operates an AI-native cloud platform that unifies three product pillars: serverless Model APIs exposing 200+ large language, image, video, audio, vision, and code models through OpenAI- and Anthropic-compatible endpoints; Agent Sandbox, a Firecracker microVM-based runtime providing kernel-isolated execution environments for autonomous coding and browser-use agents with sub-200ms startup and stateful pause/resume; and a GPU Cloud spanning dedicated GPU instances (H100, H200, B200, RTX 5090/4090), serverless GPU jobs that scale to zero, and bare-metal clusters with NVLink and 400 Gb/s RDMA interconnect. Dedicated Endpoints (Standard 98% SLA / Pro 99.5% SLA) sit alongside for enterprise production workloads.

The business model is predominantly usage-based — per-token API billing, hourly GPU rentals, per-second sandbox execution, and per-asset media generation — with a subscription layer for Dedicated Endpoints and an affiliate channel paying 10% commissions for 180 days. Pricing is positioned at up to 50% below major cloud providers (H100 SXM at $1.70/GPU/hr, B200 SXM at $4.77/GPU/hr). Go-to-market is product-led growth on a self-serve platform (sign-up includes $100 sandbox credits, no credit card required), reinforced by a Discord developer community, 50+ framework integrations (Hugging Face, LlamaIndex, LangChain, vLLM, SGLang, Claude Code, Dify, Continue, etc.), a strategic partnership with Hugging Face exposing Novita to 5M+ developers via a 'Deploy on Novita' flow, and direct enterprise sales for higher-SLA tiers. The company holds AICPA SOC 2 Type 2 certification and reports serving 350,000+ developers on the platform.

Customer concentration is mixed across named logos (Hugging Face, Quora/POE, OpenRouter, Vercel, Fish Audio, Kilo Code, Genspark, Gizmo, TiDB, Hygo, Wiz) spanning developer-tools, AI infrastructure, search, and audio verticals. The model API business is essentially a commoditized inference reseller exposed to GPU cost and competitive price compression, while the Agent Sandbox line and the Hugging Face channel partnership are the more differentiated, defensible assets. Headcount of 1–10 employees against the operational footprint claimed (1,000+ GPU H100 cluster, 24/7 inference, SOC 2 controls) implies either heavy reliance on automated/managed operations or significant under-reporting of true workforce.

Short descriptiontext

Novita AI is a San Francisco-based AI-native cloud platform that provides serverless model APIs for 200+ LLMs and generative media models, Firecracker-microVM-based Agent Sandbox runtime for autonomous AI agents, and GPU cloud infrastructure (H100, H200, B200, RTX 5090/4090), targeting developers and AI builders via a self-serve platform and a strategic Hugging Face distribution partnership.

Operating statusenum
Operating
Ownership categoryenum
Headcount rangeband
11–50
akta.pro rankint
HeadquartersSan Francisco, United States
HQ citystring
San Francisco
HQ countrystring
United States
HQ regionstring
North America
Markets served

Serves global market

Offices1 record

Each record includes

City, Country, Type, Description, Source

Keyword5 values
AI inference APIs, GPU cloud infrastructure, serverless model deployment, autonomous agent sandbox, AI model hosting
Industry3 codes
1Model Hosting, Serving & Inference Platforms
CodeHDAAACABPrimaryYes
2AI Compute Virtualization & Scheduling (GPU virtualization, cluster schedulers)
CodeHDAAAAAGPrimaryNo
3Responsible AI, Security & Privacy Platforms (Safety, Guardrails, PII)
CodeHDAEANAGPrimaryNo
NAICS code3 codes
  • Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services5182
  • Software Publishers513210
  • Computer Systems Design and Related Services54151
SIC code3 codes
  • Services-Prepackaged Software7372
  • Services-Computer Integrated Systems Design7373
  • Services-Computer Processing & Data Preparation7374
Product category
AI Cloud Infrastructure
Social media profiles2 records
GTM motion3 records

Each record includes

Type, Description, Source

Revenue model6 records
1Serverless Model API (Pay-per-token)
TypeUsage Based
Description

Customers pay per token consumed (input and output tokens) for LLM, image, audio, and video generation via serverless endpoints. No infrastructure to manage. Supports batch inference at 50% discount on input/output tokens. Revenue scales with customer usage.

novita.ai
2GPU Instance Rentals
TypeUsage Based
Description

Dedicated GPU machines (H100 SXM at $1.70/GPU/hr, B200 SXM at $4.77/GPU/hr, RTX 5090, RTX 4090) with hourly billing. Customers fully control the infrastructure with predictable performance and no shared resources.

novita.ai
3GPU Bare Metal
TypeUsage Based
Description

Reserved physical GPU clusters for large-scale inference, training, and enterprise deployments with contractual SLAs and guaranteed delivery. Billed hourly at premium rates.

novita.ai
4Dedicated Endpoints
TypeSubscription Recurring
Description

Private model endpoints with guaranteed performance and isolated resources. Priced on a monthly subscription basis with Standard (98% SLA) and Pro (99.5% SLA) tiers.

novita.ai
5Agent Sandbox
TypeUsage Based
Description

Secure isolated runtimes for AI agents. Billing per second for sandbox execution time (vCPU hours, RAM hours, storage). Available via managed NovitaClaw deployments, Sandbox Skills for OpenClaw, and Hermes-compatible runtimes.

novita.ai
6Affiliate Program
TypeAffiliate Referral
Description

Partners earn 10% commission on referral spending for the first 180 days. Commission is a cost acquisition channel for Novita.

novita.ai
Marketing channels8 records

Each record includes

Title, Type, Stage, Description, Source

Distribution channels5 records

Each record includes

Title, Type, Scope, Target buyer, Description, Source

Cost components5 values
Technology or R&D, Infrastructure, Personnel, Marketing or Sales, Operations
Pricing details8 tiers
1LLM Serverless API - usage-based per token
ModelUsage-basedBilling cadencePay-as-you-go
Notes

LLM pricing per million tokens: Deepseek V4 Pro $1.6/Mt input / $3.2/Mt output; Deepseek V4 Flash $0.14/$0.28; Qwen3.5-397B-A17B $0.6/$3.6; Gemma 4 31B $0.14/$0.4; Llama 3.1 8B $0.02/$0.05. Cache Read pricing available (e.g., Deepseek V4 Pro $0.135/Mt).

status.novita.ai
2GPU Instance rentals - hourly rate
ModelUsage-basedBilling cadencePay-as-you-go
Notes

H100 SXM (8x per node): $1.70/GPU/hr; B200 SXM: $4.77/GPU/hr; H200 SXM (8x per node, 141 GB HBM3e per GPU); RTX 5090 (8x per node, 32 GB GDDR7); RTX 4090 (8x per node, 24 GB GDDR6X).

status.novita.ai
3GPU Bare Metal - hourly rate
ModelUsage-basedBilling cadencePay-as-you-go
Notes

Bare metal GPU servers: H100 SXM at $1.70/GPU/hr (best value tier); B200 SXM at $4.77/GPU/hr (top performance tier).

status.novita.ai
4Agent Sandbox - billed per second
ModelUsage-basedBilling cadencePay-as-you-go
Notes

Agent Sandbox charges per second of execution: vCPU hours, RAM hours, and storage consumption tracked. New sign-ups receive $100 in Sandbox credits, valid 90 days.

novita.ai
5Dedicated Endpoints - monthly subscription
ModelSubscriptionBilling cadenceMonthly
Notes

Standard Dedicated Endpoints: 98% SLA. Pro Dedicated Endpoints: 99.5% SLA with higher monthly fees. Details quote-based through sales team.

novita.ai
6Image Generation - per image
ModelUnit PricingBilling cadencePay-as-you-go
Notes

Text to Image 512x512: $0.001/image; Remove Background: $0.017/image; Flux.1 Kontext Max: $0.072/image; Hunyuan Image 3: $0.1/image.

status.novita.ai
7Video Generation - per video or per second
ModelUnit PricingBilling cadencePay-as-you-go
Notes

Text to Video (32 frames, 20 steps): $0.0307/video; Kling V1.6 I2V 5s 720P: $0.27/video; Kling V3.0 Pro I2V: $0.112/s; Wan 2.5 T2V 5s 480P: $0.25/video.

status.novita.ai
8Audio/TTS - per million characters or per voice
ModelUnit PricingBilling cadencePay-as-you-go
Notes

Fish Audio TTS: $15/1M characters; MiniMax speech-02-hd: $80/1M characters; Fish Audio Voice Cloning: $0.1/voice; MiniMax Voice Cloning: $2.4/voice.

status.novita.ai
GTM typeB2B
B2B
Offering typeSoftware
Software
Brand1 of 2 records shown
1NovitaClaw
Description

Managed deployment product for AI agents

novita.ai
+1 more record
Core offering1 text field

Novita AI is an AI-native cloud platform that runs 200+ AI models (LLMs, image, audio, video, vision) through OpenAI and Anthropic-compatible APIs, scales dedicated and serverless GPU compute (H100, H200, B200, RTX 5090/4090), and provides a Firecracker microVM-based Agent Sandbox for securely executing autonomous AI agents. Revenue is generated via per-token API billing, hourly GPU rental, per-second sandbox execution, and monthly dedicated endpoint subscriptions.

Differentiator
Functional benefit
Problem solved
Quantifiable outcome1 of 7 values shown
  • Up to 50% cost savings compared to major cloud providers
+6 more records
Product overview1 text field

Novita AI is an AI-native cloud platform providing a unified infrastructure for running AI models, scaling GPU compute, and building autonomous agents. The platform consists of three core product pillars: (1) Model APIs offering serverless access to 200+ LLMs, image, video, and audio generation models; (2) Agent Sandbox providing secure isolated runtime environments using Firecracker microVMs for autonomous coding agents; and (3) GPU Cloud encompassing GPU Instances, Serverless GPU, and Bare Metal options for flexible infrastructure deployment. Additional offerings include Dedicated Endpoints for private model access with SLA guarantees, NovitaClaw agent framework, and Sandbox Skills modules. The platform is OpenAI and Anthropic API-compatible, enabling day-zero deployment of new model releases with up to 50ms time-to-first-token performance.

Product and service7 records
1Model APIs
CategoryAI Model APIs
Description

Serverless API platform providing access to 200+ AI models spanning LLMs (Deepseek V4 Pro, Qwen3.5, GLM-5.1, Kimi K2.6, Gemma 4, Llama 4, Mistral), image generation (Flux, SDXL, Hunyuan Image 3, Seedream, Qwen Image), video generation (Kling, Hunyuan Video, Wan, MiniMax, PixVerse), audio (Fish Audio TTS, MiniMax speech), and vision models through OpenAI and Anthropic-compatible endpoints, billed per token with no infrastructure management required. Targets developers building AI-powered applications.

2Agent Sandbox
CategoryAgent Runtime Infrastructure
Description

Secure runtime infrastructure using Firecracker microVMs that isolates autonomous AI agent systems during execution with individual kernels and isolated memory boundaries, supporting sub-200ms startup times, scaling to thousands of parallel microVMs, and stateful pause/resume for long-running workflows. Designed for coding agents, API-calling agents, and web-browsing agents; available through managed NovitaClaw deployments, Sandbox Skills for OpenClaw, and Hermes-compatible runtimes.

3GPU Instance
CategoryGPU Cloud Infrastructure
Description

Full-control dedicated GPU machines available in seconds with NVIDIA H100, H200, A100, L4, RTX 4090, and RTX 5090 configurations. Offers predictable performance with isolated resources, no shared infrastructure, and supports deploying models, running inference, and training from scratch. Targets AI/ML teams needing dedicated GPU compute with hourly billing.

4Serverless GPU
CategoryGPU Cloud Infrastructure
Description

On-demand GPU job execution with automatic resource allocation, auto-scaling to zero when idle, and pay-only-for-execution billing model. No instances to provision or idle compute to manage, making it suitable for intermittent or bursty inference and batch workloads.

5GPU Bare Metal
CategoryGPU Cloud Infrastructure
Description

Dedicated physical GPU clusters (H100 SXM, B200 SXM, H200 SXM, RTX 5090, RTX 4090) for large-scale inference, training runs, and enterprise deployments requiring maximum performance with zero virtualization overhead. Billed hourly at premium rates (H100 SXM at $1.70/GPU/hr, B200 SXM at $4.77/GPU/hr) with 1,000+ GPU linear scaling on H100 SXM clusters.

6Dedicated Endpoints
CategoryManaged Inference Endpoints
Description

Private model endpoints with guaranteed performance and isolated resources ensuring consistent latency at any throughput without noisy-neighbor issues. Available in Standard (98% SLA) and Pro (99.5% SLA) tiers on a monthly subscription basis, designed for production AI deployments requiring SLA-backed reliability.

7GPU Cloud
CategoryGPU Cloud Infrastructure
Description

Unified GPU infrastructure platform encompassing GPU Instances, Serverless GPU, and Bare Metal offerings, providing flexible deployment options from fully managed serverless to dedicated physical hardware. Includes multi-node GPU clusters with NVLink 4th Gen, GPUDirect RDMA, and 400 Gb/s RDMA networking (e.g., CLUSTER-01 with 6x H200 nodes, CLUSTER-02 with 6x H100 nodes) for different workload requirements.

Scale indicator12 records

Each record includes

Type, Value, Description, Source

Partnership7 partners
Strategic tierCoreTypeStrategic or Co-development PartnerAnnounced on2026-04-14
Description

Novita AI and Hugging Face announced a strategic partnership enabling over five million developers on the Hugging Face platform to instantly deploy AI models as production-ready APIs without infrastructure setup. The partnership introduces a seamless 'Deploy on Novita' experience removing operational overhead. Novita AI was a day 0 launch partner for Google's Gemma 4 model, and is also available directly on Hugging Face's platform.

Strategic tierMinorTypeChannel Partner/ Reseller/ Distributor
Description

Novita models are available on POE (Quora's AI platform), enabling POE's user base to access Novita's inference infrastructure.

Strategic tierCoreTypeTechnology or Integration
Description

Novita AI partnered with vLLM to advance AI inference, optimizing the vLLM inference engine on Novita's GPU infrastructure for better performance and cost efficiency.

Strategic tierCoreTypeTechnology or Integration
Description

Integration with SGLang (Structured Generation Language) framework, enabling developers to build and deploy AI applications using SGLang on Novita's GPU cloud.

Strategic tierCoreTypeTechnology or Integration
Description

Official Novita AI integration with LlamaIndex, enabling RAG (Retrieval Augmented Generation) workflows using Novita's model APIs.

Strategic tierCoreTypeTechnology or Integration
Description

Official integration with LangChain framework, allowing LangChain developers to easily connect to Novita's model APIs and GPU infrastructure.

Strategic tierCoreTypeTechnology or Integration
Description

Claude Code integration guide published by Novita, enabling developers to use Claude Code with Novita's infrastructure for AI-powered coding workflows.

Recent move6 records

Each record includes

Date, Type, Title, Description, Source

Expansion highlight6 records

Each record includes

Type, Description

Peers10 records
TypeBroad incumbent
Description

Large-scale GPU cloud provider originally built for crypto/HPC workloads, now serving major AI labs with H100/B200 clusters and bare metal. Competes with Novita's Bare Metal and GPU Cluster products but operates at substantially larger scale with enterprise-only GTM.

TypeDirect peer
Description

Cloud API for running open-source ML models (LLMs, image, video, audio) on demand with per-second billing. Directly comparable to Novita's serverless multi-modal model API and similar developer-led GTM motion.

TypeDirect peer
Description

GPU cloud offering on-demand instances and serverless endpoints for AI training and inference at consumer-friendly hourly rates. Directly comparable to Novita's GPU Instance and Serverless GPU products for individual developers and small teams.

TypeOthers
Description

While Novita's largest distribution partner and customer (not direct competitor), Hugging Face Inference Endpoints offers competing dedicated inference infrastructure. Critical ecosystem anchor for Novita's developer acquisition and a potential competitor in the inference layer.

TypeDirect peer
Description

Commercial platform behind the open-source Ray distributed computing framework, offering managed compute for AI/ML workloads including LLM serving. Comparable to Novita's GPU Cloud and vLLM-optimized inference stack for production AI deployments.

TypeBroad incumbent
Description

GPU cloud and on-prem AI infrastructure provider offering H100 clusters, GPU cloud instances, and model API services. Overlaps with Novita's GPU Cloud and Bare Metal offerings but with a heavier on-prem and training orientation.

TypeBroad incumbent
Description

Vertically-integrated GPU cloud built around stranded energy, serving AI training and inference workloads with H100 clusters. Comparable to Novita's Bare Metal and GPU Cluster offerings; co-listed with Novita on Ramp's trending AI infrastructure list, signaling similar customer overlap.

TypeDirect peer
Description

Serverless cloud platform for running AI/ML workloads with Python-based infrastructure-as-code and per-second GPU billing. Overlaps with Novita's Serverless GPU and Agent Sandbox offerings for developers running code-driven AI pipelines.

TypeDirect peer
Description

Production inference platform for open and proprietary LLMs with serverless APIs and dedicated deployments, focused on low-latency, cost-efficient serving. Competes head-to-head with Novita on the same token-priced model API product.

TypeDirect peer
Description

Open-source-focused AI inference cloud offering serverless model APIs and dedicated GPU instances across 200+ open models. Closest direct competitor to Novita's model API and GPU cloud pillars, with comparable OpenAI-compatible endpoints and similar price-performance positioning.

Market position
Strengths5 records

Each record includes

Headline, Details, Source

Weaknesses5 records

Each record includes

Headline, Details, Source

Competitive moat6 records

Each record includes

Type, Details

Key risks6 records

Each record includes

Headline, Details, Source

Key highlights7 records

Each record includes

Headline, Details, Source

Customer concentration

Classification, Details

Named customers9 records

Each record includes

Name, Industry, Type, Use case, Source, UUID

Segment4 records

Each record includes

Title, Type, Primary, Description, Pain point addressed, Use case, Source

Ideal customer profile5 records

Each record includes

Profile, Firmographic size, Sales motion, Sales cycle length, Buying structure, Purchase trigger, Buyer persona, Geography, Industry vertical, Primary use case, Description, Pain points, Evidence proof points, Target buyer

Technology focused
Yes
API detail
Has APIbool
Yes

Docs URL, Description

Integration16 records

Each record includes

Title, Type, Description, Source

AI capability15 records

Each record includes

Type, Description, Source

AI maturity
App detail

Has app

Feature5 records

Each record includes

Title, Differentiator, Description, Source

Core technology
Revenue estimate
Valuation estimate
Number of profiles
No data
Compliance1 record

Each record includes

Name, Class, Description

Funding overview

Funding stage, Last funding date, Total funding USD

Funding rounds

Each record includes

Round, Amount USD, Date, Pre money valuation, Total investors, Investors, News

Investors

Each record includes

Name, Type, Date of entry, Rounds participated, Website

Funding detail is available on the Subscription and Enterprise plan.Contact sales →

M&A

Each record includes

Name, Acquisition type, Announced date, Completed date, Status, Website, News

Investment

Each record includes

Name, Round, Announced date, Lead investor, Website, News

M&A and investment is available on the Subscription and Enterprise plan.Contact sales →

Novita AI

AI Cloud Infrastructurenovita.ai

Novita AI is a San Francisco-based AI-native cloud platform that provides serverless model APIs for 200+ LLMs and generative media models, Firecracker-microVM-based Agent Sandbox runtime for autonomous AI agents, and GPU cloud infrastructure (H100, H200, B200, RTX 5090/4090), targeting developers and AI builders via a self-serve platform and a strategic Hugging Face distribution partnership.

What Novita AI does

Novita AI (San Francisco, founded 2024) operates an AI-native cloud platform that unifies three product pillars: serverless Model APIs exposing 200+ large language, image, video, audio, vision, and code models through OpenAI- and Anthropic-compatible endpoints; Agent Sandbox, a Firecracker microVM-based runtime providing kernel-isolated execution environments for autonomous coding and browser-use agents with sub-200ms startup and stateful pause/resume; and a GPU Cloud spanning dedicated GPU instances (H100, H200, B200, RTX 5090/4090), serverless GPU jobs that scale to zero, and bare-metal clusters with NVLink and 400 Gb/s RDMA interconnect. Dedicated Endpoints (Standard 98% SLA / Pro 99.5% SLA) sit alongside for enterprise production workloads.

The business model is predominantly usage-based — per-token API billing, hourly GPU rentals, per-second sandbox execution, and per-asset media generation — with a subscription layer for Dedicated Endpoints and an affiliate channel paying 10% commissions for 180 days. Pricing is positioned at up to 50% below major cloud providers (H100 SXM at $1.70/GPU/hr, B200 SXM at $4.77/GPU/hr). Go-to-market is product-led growth on a self-serve platform (sign-up includes $100 sandbox credits, no credit card required), reinforced by a Discord developer community, 50+ framework integrations (Hugging Face, LlamaIndex, LangChain, vLLM, SGLang, Claude Code, Dify, Continue, etc.), a strategic partnership with Hugging Face exposing Novita to 5M+ developers via a 'Deploy on Novita' flow, and direct enterprise sales for higher-SLA tiers. The company holds AICPA SOC 2 Type 2 certification and reports serving 350,000+ developers on the platform.

Customer concentration is mixed across named logos (Hugging Face, Quora/POE, OpenRouter, Vercel, Fish Audio, Kilo Code, Genspark, Gizmo, TiDB, Hygo, Wiz) spanning developer-tools, AI infrastructure, search, and audio verticals. The model API business is essentially a commoditized inference reseller exposed to GPU cost and competitive price compression, while the Agent Sandbox line and the Hugging Face channel partnership are the more differentiated, defensible assets. Headcount of 1–10 employees against the operational footprint claimed (1,000+ GPU H100 cluster, 24/7 inference, SOC 2 controls) implies either heavy reliance on automated/managed operations or significant under-reporting of true workforce.

Novita AI firmographics

Firmographics
Name
Novita AI
Legal name
Novita AI
Website
https://novita.ai
Company type
Private
Founded year
2024
Operating status
Operating
Headcount range
11–50 employees
Short description
Novita AI is a San Francisco-based AI-native cloud platform that provides serverless model APIs for 200+ LLMs and generative media models, Firecracker-microVM-based Agent Sandbox runtime for autonomous AI agents, and GPU cloud infrastructure (H100, H200, B200, RTX 5090/4090), targeting developers and AI builders via a self-serve platform and a strategic Hugging Face distribution partnership.
Ownership category
akta.pro rank

Novita AI industry classification

Industry
Product category
AI Cloud Infrastructure
NAICS
Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (5182), Software Publishers (513210), Computer Systems Design and Related Services (54151)
SIC
Services-Prepackaged Software (7372), Services-Computer Integrated Systems Design (7373), Services-Computer Processing & Data Preparation (7374)
akta.pro primary industry
Model Hosting, Serving & Inference Platforms (HDAAACAB)
akta.pro secondary industries
AI Compute Virtualization & Scheduling (GPU virtualization, cluster schedulers) (HDAAAAAG), Responsible AI, Security & Privacy Platforms (Safety, Guardrails, PII) (HDAEANAG)

Keywords

  • AI inference APIs
  • GPU cloud infrastructure
  • Serverless model deployment
  • Autonomous agent sandbox
  • AI model hosting

Where Novita AI is headquartered

Location

Headquarters

HQ city
San Francisco
HQ country
United States
HQ region
North America

Offices1 record

Markets served

Novita AI business model

Business model
GTM type
B2B
Offering type
Software
Cost components
Technology or R&D, Infrastructure, Personnel, Marketing or Sales, Operations

Revenue model

  1. Serverless Model API (Pay-per-token): Customers pay per token consumed (input and output tokens) for LLM, image, audio, and video generation via serverless endpoints. No infrastructure to manage. Supports batch inference at 50% discount on input/output tokens. Revenue scales with customer usage.
  2. GPU Instance Rentals: Dedicated GPU machines (H100 SXM at $1.70/GPU/hr, B200 SXM at $4.77/GPU/hr, RTX 5090, RTX 4090) with hourly billing. Customers fully control the infrastructure with predictable performance and no shared resources.
  3. GPU Bare Metal: Reserved physical GPU clusters for large-scale inference, training, and enterprise deployments with contractual SLAs and guaranteed delivery. Billed hourly at premium rates.
  4. Dedicated Endpoints: Private model endpoints with guaranteed performance and isolated resources. Priced on a monthly subscription basis with Standard (98% SLA) and Pro (99.5% SLA) tiers.
  5. Agent Sandbox: Secure isolated runtimes for AI agents. Billing per second for sandbox execution time (vCPU hours, RAM hours, storage). Available via managed NovitaClaw deployments, Sandbox Skills for OpenClaw, and Hermes-compatible runtimes.
  6. Affiliate Program: Partners earn 10% commission on referral spending for the first 180 days. Commission is a cost acquisition channel for Novita.

Pricing tiers

ModelBillingPrice
Usage-basedPay-as-you-goLLM Serverless API - usage-based per token
Usage-basedPay-as-you-goGPU Instance rentals - hourly rate
Usage-basedPay-as-you-goGPU Bare Metal - hourly rate
Usage-basedPay-as-you-goAgent Sandbox - billed per second
SubscriptionMonthlyDedicated Endpoints - monthly subscription
Unit PricingPay-as-you-goImage Generation - per image
Unit PricingPay-as-you-goVideo Generation - per video or per second
Unit PricingPay-as-you-goAudio/TTS - per million characters or per voice

Go-to-market motion3 records

Distribution channels5 records

Marketing channels8 records

Novita AI product offering

Product offering

Core offering

Novita AI is an AI-native cloud platform that runs 200+ AI models (LLMs, image, audio, video, vision) through OpenAI and Anthropic-compatible APIs, scales dedicated and serverless GPU compute (H100, H200, B200, RTX 5090/4090), and provides a Firecracker microVM-based Agent Sandbox for securely executing autonomous AI agents. Revenue is generated via per-token API billing, hourly GPU rental, per-second sandbox execution, and monthly dedicated endpoint subscriptions.

Product overview

Novita AI is an AI-native cloud platform providing a unified infrastructure for running AI models, scaling GPU compute, and building autonomous agents. The platform consists of three core product pillars: (1) Model APIs offering serverless access to 200+ LLMs, image, video, and audio generation models; (2) Agent Sandbox providing secure isolated runtime environments using Firecracker microVMs for autonomous coding agents; and (3) GPU Cloud encompassing GPU Instances, Serverless GPU, and Bare Metal options for flexible infrastructure deployment. Additional offerings include Dedicated Endpoints for private model access with SLA guarantees, NovitaClaw agent framework, and Sandbox Skills modules. The platform is OpenAI and Anthropic API-compatible, enabling day-zero deployment of new model releases with up to 50ms time-to-first-token performance.

Differentiator

Problem solved

Functional benefit

Brands

  • NovitaClaw: Managed deployment product for AI agents
  • Novita Sandbox

Products and services

  • Model APIs Serverless API platform providing access to 200+ AI models spanning LLMs (Deepseek V4 Pro, Qwen3.5, GLM-5.1, Kimi K2.6, Gemma 4, Llama 4, Mistral), image generation (Flux, SDXL, Hunyuan Image 3, Seedream, Qwen Image), video generation (Kling, Hunyuan Video, Wan, MiniMax, PixVerse), audio (Fish Audio TTS, MiniMax speech), and vision models through OpenAI and Anthropic-compatible endpoints, billed per token with no infrastructure management required. Targets developers building AI-powered applications.
  • Agent Sandbox Secure runtime infrastructure using Firecracker microVMs that isolates autonomous AI agent systems during execution with individual kernels and isolated memory boundaries, supporting sub-200ms startup times, scaling to thousands of parallel microVMs, and stateful pause/resume for long-running workflows. Designed for coding agents, API-calling agents, and web-browsing agents; available through managed NovitaClaw deployments, Sandbox Skills for OpenClaw, and Hermes-compatible runtimes.
  • GPU Instance Full-control dedicated GPU machines available in seconds with NVIDIA H100, H200, A100, L4, RTX 4090, and RTX 5090 configurations. Offers predictable performance with isolated resources, no shared infrastructure, and supports deploying models, running inference, and training from scratch. Targets AI/ML teams needing dedicated GPU compute with hourly billing.
  • Serverless GPU On-demand GPU job execution with automatic resource allocation, auto-scaling to zero when idle, and pay-only-for-execution billing model. No instances to provision or idle compute to manage, making it suitable for intermittent or bursty inference and batch workloads.
  • GPU Bare Metal Dedicated physical GPU clusters (H100 SXM, B200 SXM, H200 SXM, RTX 5090, RTX 4090) for large-scale inference, training runs, and enterprise deployments requiring maximum performance with zero virtualization overhead. Billed hourly at premium rates (H100 SXM at $1.70/GPU/hr, B200 SXM at $4.77/GPU/hr) with 1,000+ GPU linear scaling on H100 SXM clusters.
  • Dedicated Endpoints Private model endpoints with guaranteed performance and isolated resources ensuring consistent latency at any throughput without noisy-neighbor issues. Available in Standard (98% SLA) and Pro (99.5% SLA) tiers on a monthly subscription basis, designed for production AI deployments requiring SLA-backed reliability.
  • GPU Cloud Unified GPU infrastructure platform encompassing GPU Instances, Serverless GPU, and Bare Metal offerings, providing flexible deployment options from fully managed serverless to dedicated physical hardware. Includes multi-node GPU clusters with NVLink 4th Gen, GPUDirect RDMA, and 400 Gb/s RDMA networking (e.g., CLUSTER-01 with 6x H200 nodes, CLUSTER-02 with 6x H100 nodes) for different workload requirements.

Quantifiable outcome

  • Up to 50% cost savings compared to major cloud providers
  • +6 more outcomes

Companies that use Novita AI

Customer profile

Named customers9 records

Segments4 records

Ideal customer profiles5 records

Novita AI technology and API

Technology

Technology focussed Yes

API detail

Has API
Yes
API docs
API detail

Core technology

AI maturity

App detail

Integration16 records

AI capability15 records

Feature5 records

Novita AI partnerships and signals

Strategic signal

Partnerships

Seven partnerships are on record, tiered core and minor.

  • Hugging FacecoreStrategic or Co-development Partner · 14 April 2026Novita AI and Hugging Face announced a strategic partnership enabling over five million developers on the Hugging Face platform to instantly deploy AI models as production-ready APIs without infrastructure setup. The partnership introduces a seamless 'Deploy on Novita' experience removing operational overhead. Novita AI was a day 0 launch partner for Google's Gemma 4 model, and is also available directly on Hugging Face's platform.
  • POEminorChannel Partner/ Reseller/ DistributorNovita models are available on POE (Quora's AI platform), enabling POE's user base to access Novita's inference infrastructure.
  • vLLMcoreTechnology or IntegrationNovita AI partnered with vLLM to advance AI inference, optimizing the vLLM inference engine on Novita's GPU infrastructure for better performance and cost efficiency.
  • SGLangcoreTechnology or IntegrationIntegration with SGLang (Structured Generation Language) framework, enabling developers to build and deploy AI applications using SGLang on Novita's GPU cloud.
  • LlamaIndexcoreTechnology or IntegrationOfficial Novita AI integration with LlamaIndex, enabling RAG (Retrieval Augmented Generation) workflows using Novita's model APIs.
  • LangchaincoreTechnology or IntegrationOfficial integration with LangChain framework, allowing LangChain developers to easily connect to Novita's model APIs and GPU infrastructure.
  • Claude (Anthropic)coreTechnology or IntegrationClaude Code integration guide published by Novita, enabling developers to use Claude Code with Novita's infrastructure for AI-powered coding workflows.

Scale indicators12 records

Recent moves6 records

Expansion highlights6 records

Novita AI competitors and assessment

Company assessment

Broad incumbents

  • CoreWeave: Large-scale GPU cloud provider originally built for crypto/HPC workloads, now serving major AI labs with H100/B200 clusters and bare metal. Competes with Novita's Bare Metal and GPU Cluster products but operates at substantially larger scale with enterprise-only GTM.
  • Lambda Labs: GPU cloud and on-prem AI infrastructure provider offering H100 clusters, GPU cloud instances, and model API services. Overlaps with Novita's GPU Cloud and Bare Metal offerings but with a heavier on-prem and training orientation.
  • Crusoe: Vertically-integrated GPU cloud built around stranded energy, serving AI training and inference workloads with H100 clusters. Comparable to Novita's Bare Metal and GPU Cluster offerings; co-listed with Novita on Ramp's trending AI infrastructure list, signaling similar customer overlap.

Direct peers

  • Replicate: Cloud API for running open-source ML models (LLMs, image, video, audio) on demand with per-second billing. Directly comparable to Novita's serverless multi-modal model API and similar developer-led GTM motion.
  • RunPod: GPU cloud offering on-demand instances and serverless endpoints for AI training and inference at consumer-friendly hourly rates. Directly comparable to Novita's GPU Instance and Serverless GPU products for individual developers and small teams.
  • Anyscale: Commercial platform behind the open-source Ray distributed computing framework, offering managed compute for AI/ML workloads including LLM serving. Comparable to Novita's GPU Cloud and vLLM-optimized inference stack for production AI deployments.
  • Modal Labs: Serverless cloud platform for running AI/ML workloads with Python-based infrastructure-as-code and per-second GPU billing. Overlaps with Novita's Serverless GPU and Agent Sandbox offerings for developers running code-driven AI pipelines.
  • Fireworks AI: Production inference platform for open and proprietary LLMs with serverless APIs and dedicated deployments, focused on low-latency, cost-efficient serving. Competes head-to-head with Novita on the same token-priced model API product.
  • Together AI: Open-source-focused AI inference cloud offering serverless model APIs and dedicated GPU instances across 200+ open models. Closest direct competitor to Novita's model API and GPU cloud pillars, with comparable OpenAI-compatible endpoints and similar price-performance positioning.

Others

  • Hugging Face: While Novita's largest distribution partner and customer (not direct competitor), Hugging Face Inference Endpoints offers competing dedicated inference infrastructure. Critical ecosystem anchor for Novita's developer acquisition and a potential competitor in the inference layer.

Market position

Strengths5 records

Weaknesses5 records

Competitive moat6 records

Key risks6 records

Key highlights7 records

Customer concentration

Novita AI social profiles

Digital presence

Novita AI compliance and trust

Trust signal

Compliance1 record

Novita AI financial estimates

Financial estimate

Revenue estimate

Valuation estimate

Novita AI leadership team

Management profile

Number of profiles

Novita AI funding detail

Funding detail

Funding overview

Funding rounds

Investors

Funding detail is available on the Subscription and Enterprise plan.Contact sales →

Novita AI M&A and investment

M&A and investment

M&A

Investments

M&A and investment is available on the Subscription and Enterprise plan.Contact sales →

Frequently asked questions about Novita AI

What does Novita AI do?

Novita AI is an AI-native cloud platform that runs 200+ AI models (LLMs, image, audio, video, vision) through OpenAI and Anthropic-compatible APIs, scales dedicated and serverless GPU compute (H100, H200, B200, RTX 5090/4090), and provides a Firecracker microVM-based Agent Sandbox for securely executing autonomous AI agents. Revenue is generated via per-token API billing, hourly GPU rental, per-second sandbox execution, and monthly dedicated endpoint subscriptions.

Is Novita AI a public or private company?

Novita AI is a private company. It is classified as unknown and is currently operating.

When was Novita AI founded?

Novita AI was founded in 2024. It employs 11 to 50 people.

Where is Novita AI based?

Novita AI is headquartered in San Francisco, United States, in the North America region.

How does Novita AI make money?

Six revenue lines are on record. Serverless Model API (Pay-per-token) is the primary driver. The others are GPU Instance Rentals, GPU Bare Metal, dedicated Endpoints, agent Sandbox and affiliate Program.

Who are Novita AI's main competitors?

Broad incumbents on record are CoreWeave, Lambda Labs and Crusoe. Direct peers are Replicate, RunPod, Anyscale, Modal Labs, Fireworks AI and Together AI. Hugging Face is listed as an others.

Does Novita AI have an API?

Yes. Novita AI provides a comprehensive API platform offering OpenAI-compatible and Anthropic-compatible endpoints for LLM inference, image generation, video generation, audio processing, and model fine-tuning. The platform supports serverless and dedicated endpoint deployments with authentication via Bearer token. API documentation is available at https://novita.ai/docs with endpoints documented for GPU instances, model APIs, and sandbox management. MCP (Model Context Protocol) server is available as indicated by the Agent Discovery Metadata listing /.well-known/mcp/server-card.json. Developer documentation is at novita.ai/docs.

What industry is Novita AI in?

Novita AI's product category is AI Cloud Infrastructure. Its primary akta.pro industry code is HDAAACAB, Model Hosting, Serving & Inference Platforms, with a secondary code of HDAAAAAG, AI Compute Virtualization & Scheduling (GPU virtualization, cluster schedulers). Its NAICS code is 5182 and its SIC code is 7372.

Unlock the full company data

50 free credits on sign-up, no credit card required.

Contact sales
Live signals
Pulse 2.0Novita AI Launches Sandbox To Secure Autonomous Agent Systems At Enterprise ScaleNovita AI, a San Francisco-based AI and agent cloud platform, has launched Novita Sandbox, a secure runtime infrastructure layer that isolates autonomous agent systems during execution using dedicated Firecracker microVMs with individual kernels and isolated memory boundaries. The product targets security risks introduced by autonomous agents that execute code, call APIs, browse the web, and interact with live environments, delivering sub-200 millisecond startup times while scaling to thousands of parallel microVMs for production-scale deployments. Novita Sandbox is available immediately through managed NovitaClaw deployments, Sandbox Skills for OpenClaw, Hermes-compatible runtimes, and persistent node configurations.AijournNovita AI Launches Sandbox to Secure OpenClaw, Hermes Agent, and Autonomous SystemsNovita AI announced the launch of Novita Sandbox, a security product designed to provide system-level isolation for autonomous AI agent systems using Firecracker microVMs. The product offers sub-200ms startup times, scalable parallel execution, and stateful pause/resume capabilities for long-running workflows. It is designed to address systemic security risks in agent frameworks like OpenClaw and Hermes Agent, including credential leakage, prompt injection, and cross-agent interference.PR NewswireNovita AI Launches Sandbox to Secure OpenClaw, Hermes Agent, and Autonomous SystemsNovita AI has launched Novita Sandbox, a security infrastructure product that uses Firecracker microVMs to provide system-level isolation for autonomous AI agent systems, enabling safe deployment at scale without exposing sensitive environments. The product delivers sub-200ms average startup time and supports frameworks including OpenClaw and Hermes Agent, with stateful execution capabilities allowing environments to be paused and resumed while maintaining full runtime state. The sandbox is available through managed NovitaClaw deployments, Sandbox Skills, Hermes-compatible runtimes, and persistent node configurations for long-running workflows.PR NewswireNovita AI Ranked as the Best Performing & Reliable Inference LayerIndependent AI benchmarking firm Artificial Analysis ranked Novita AI as the top inference provider in its GPT-OSS-120B assessment, placing the company first in scientific reasoning accuracy with a GPQA Diamond score of 79.0% and at the level of top providers for advanced mathematics at 93.3% on AIME 2025. Novita AI's platform hosts over 120 large language models through a single OpenAI and Anthropic-compatible API, offering day-zero availability for new model releases. The company counts Hugging Face, Quora, OpenRouter, Vercel, Kilo Code, and Genspark among its customer base.PR NewswireNovita AI Joins Hugging Face as Official Inference PartnerNovita AI and Hugging Face announced a strategic partnership on April 14, 2026, enabling over five million developers on Hugging Face's platform to instantly deploy AI models as production-ready APIs without infrastructure setup. Novita AI, which was a day 0 launch partner for Google's Gemma 4 model, offers inference with time-to-first-token as low as 50ms and potential cost savings of up to 50% compared to traditional endpoints. The partnership introduces a seamless "Deploy on Novita" experience that removes operational overhead for development teams.AijournNovita AI named a top AI infrastructure vendor on RampNovita AI was recognized as a top AI infrastructure vendor on Ramp's February 2026 Top Software Vendors list, a ranking based on actual business expenditure data from more than 50,000 companies using Ramp's corporate card and bill pay platform. The recognition places Novita alongside Cerebras, Runware, Clarifai, Crusoe, and Modal as trending AI infrastructure companies, with the category being renamed from "AI infrastructure" to "Agent Hosting and Serving." Novita AI currently serves over 350,000 developers who utilize more than one trillion tokens per day to build and run AI agents.PR NewswireNovita AI named a top AI infrastructure vendor on RampNovita AI has been recognized as one of the top AI infrastructure vendors on Ramp's February 2026 Top Software Vendors list, based on actual business expenditure data from more than 50,000 companies using Ramp's corporate card and bill pay platform. The recognition places Novita alongside other trending AI infrastructure companies including Cerebras, Runware, Clarifai, Crusoe, and Modal. Ramp also renamed the AI infrastructure category to "Agent Hosting and Serving" to reflect the growing prominence of AI agents, with Novita highlighting its Agent Sandbox feature that provides secure, isolated environments for running agents.MorningstarNovita AI named a top AI infrastructure vendor on RampNovita AI was recognized as a top AI infrastructure vendor on Ramp's February 2026 Top Software Vendors list, a ranking based on actual business expenditure data from more than 50,000 companies using Ramp's corporate card and bill pay platform. The recognition places Novita AI alongside other trending AI infrastructure companies including Cerebras, Runware, Clarifai, Crusoe, and Modal. The article also highlights Novita AI's Agent Sandbox product, which enables users to run AI agents in a fully isolated environment.