Radium
Radium provides enterprise AI inference through a bare-metal infrastructure platform, offering OpenAI-compatible API access to three model tiers (Hal, Clarke, Tycho) at approximately half the cost of frontier providers, targeting enterprise teams seeking to reduce AI infrastructure spend.
- Company typePrivate
- Founded2020
- HeadquartersSan Francisco, United States
- Headcount11–50
- GTM typeB2B
- OfferingSoftware
What Radium does
Radium (legal entity: Radium Enterprise AI) is an enterprise AI inference provider incorporated in Ontario, Canada, with an operational base in San Francisco. The company offers a swap-in, OpenAI- and Anthropic-compatible inference platform targeting enterprise teams that are already integrated with frontier-model APIs and are seeking to reduce AI infrastructure spend. Radium's go-to-market positions the service as a cost- and performance-optimized alternative for production AI workloads, claiming approximately 50% lower cost than comparable OpenAI and Anthropic endpoints at equivalent capability.
The platform is anchored by a proprietary bare-metal infrastructure stack designed from the kernel level up for AI workloads. Core technical components include a zero-virtualization kernel optimized for GPU locality and predictable performance (marketed as ~0 jitter), an RL-based intelligent orchestration engine that performs workload-aware scheduling across NUMA, SXM, PCIe, and fabric topologies, and a custom dual-plane network fabric separating InfiniBand AI traffic from Ethernet control/storage. Radium delivers three model tiers — Hal 1.0 (maximum capability, comparable to Opus 4.7 and GPT-5.5), Clarke 1.0 (balanced performance, comparable to Sonnet 4.6 and GPT-5.4), and Tycho 1.0 (high-efficiency scale, comparable to Haiku 4.5 and GPT-5.4 Mini) — all built on open-source foundations and exposed through an OpenAI-compatible API contract that allows single-endpoint migration.
Radium generates revenue through usage-based, per-token API pricing across its three model tiers, with input and output tokens billed separately. Distribution combines a self-serve signup platform (deploy.radium.cloud) with direct enterprise sales and an OEM/embedded channel demonstrated through the Realbotix partnership. Named reference customers include Square (R&D prototyping), EQTY Lab (climate model training, COP28), Realbotix (humanoid robotics inference), and Alexi (legal-tech retrieval models), spanning financial services, climate, robotics, and legal verticals. Pricing examples indicate approximately $1,600/month for 320M tokens on the highest tier versus $3,200/month for equivalent OpenAI or Anthropic endpoints, with optional prompt caching providing up to 90% savings on repeated prompts.
Radium firmographics
Firmographics- Name
- Radium
- Legal name
- Radium Enterprise AI
- Website
- https://radium.cloud
- Company type
- Private
- Founded year
- 2020
- Operating status
- Operating
- Headcount range
- 11–50 employees
- Short description
- Radium provides enterprise AI inference through a bare-metal infrastructure platform, offering OpenAI-compatible API access to three model tiers (Hal, Clarke, Tycho) at approximately half the cost of frontier providers, targeting enterprise teams seeking to reduce AI infrastructure spend.
- Ownership category
- akta.pro rank
Where Radium is headquartered
LocationHeadquarters
- HQ city
- San Francisco
- HQ country
- United States
- HQ region
- North America
Offices1 record
Markets served
Radium business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Infrastructure, Technology or R&D, Personnel, Operations, Marketing or Sales
Revenue model
- API Token-based Inference: Radium generates revenue through consumption-based API access to AI model inference endpoints. Customers pay per token (input and output tokens separately) based on which model tier they use (Hal, Clarke, or Tycho). The pricing is structured to be approximately 50% lower than comparable OpenAI and Anthropic endpoints. Additional cost optimization comes from prompt caching which can save up to 90% on repeated prompts.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Usage-based | Pay-as-you-go | Hal 1.0 - Maximum capability model for complex reasoning and agent workflows |
| Usage-based | Pay-as-you-go | Hal 1.0 - Maximum capability model for complex reasoning and agent workflows |
| Usage-based | Pay-as-you-go | Clarke 1.0 - Balanced performance model for general-purpose enterprise workloads |
| Usage-based | Pay-as-you-go | Tycho 1.0 - High-efficiency scale model for high-volume, cost-sensitive workloads |
Go-to-market motion1 record
Distribution channels3 records
Marketing channels5 records
Radium product offering
Product offeringCore offering
Radium is an OpenAI/Anthropic-compatible AI inference platform that delivers three open-source-based model tiers (Hal 1.0, Clarke 1.0, Tycho 1.0) over a serverless API. Customers access the models by swapping the endpoint URL—a single line of code—and pay per input and output token at approximately 50% of comparable frontier provider pricing, with optional prompt caching that reduces costs on repeated prompts by up to 90%.
Product overview
Radium is an enterprise AI inference platform offering three distinct model tiers built on open-source foundations and delivered through proprietary bare-metal infrastructure. The core portfolio includes: (1) Hal 1.0 for maximum capability complex reasoning and agentic workflows, (2) Clarke 1.0 for balanced production workloads, and (3) Tycho 1.0 for high-efficiency, high-volume tasks. All models are delivered via a unified serverless inference platform with OpenAI/Claude-compatible API, supported by Radium's zero-virtualization bare-metal infrastructure that eliminates abstraction layers between software and hardware for optimized GPU utilization and reduced latency.
Differentiator
Problem solved
Functional benefit
Brands
- Hal 1.0: Maximum capability model tier for complex reasoning, agent workflows, and software tasks.
- Clarke 1.0
- Tycho 1.0
Products and services
- Hal 1.0 Hal 1.0 is Radium's maximum capability AI model tier for complex reasoning, agentic workflows, code generation, architecture-level planning, and multi-file software tasks. It is comparable to Anthropic Opus 4.7 and OpenAI GPT-5.5 and is priced at $2.25 per million input tokens and $11.50 per million output tokens on a pay-as-you-go basis, accessed via Radium's OpenAI/Claude-compatible API.
- Clarke 1.0 Clarke 1.0 is Radium's balanced performance AI model tier for general-purpose enterprise workloads, including document reasoning, structured outputs (JSON), code generation and debugging, RAG, and enterprise copilot applications. Comparable to Anthropic Sonnet 4.6 and OpenAI GPT-5.4, priced at $1.50 per million input tokens and $7.00 per million output tokens on pay-as-you-go.
- Tycho 1.0 Tycho 1.0 is Radium's high-efficiency AI model tier for high-volume, cost-sensitive workloads such as classification, extraction, summarization, ticket triage, intent detection, document labeling, and lightweight agent execution. Comparable to Anthropic Haiku 4.5 and OpenAI GPT-5.4 Mini, priced at $0.50 per million input tokens and $2.25 per million output tokens on pay-as-you-go.
Quantifiable outcome
- 2x lower inference latency vs AWS (Carnegie Mellon University benchmark)
- +4 more outcomes
Companies that use Radium
Customer profileNamed customers4 records
Segments3 records
Ideal customer profiles2 records
Radium technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
AI capability9 records
Feature4 records
Radium partnerships and signals
Strategic signalPartnerships
Four partnerships are on record, tiered core.
- RealbotixcoreRealbotix Corp. (TSX-V: XBOT) announced a strategic collaboration with Radium to power real-time AI companions using Radium's serverless inference platform, addressing critical latency challenges in consumer robotics. The partnership enables lightning-fast conversations and GPU auto-scaling for Realbotix's humanoid robots built for entertainment, customer service, and personal well-being applications.
- Stanford UniversitycoreStanford University's Center for Research on Foundation Models conducted independent benchmarks comparing Radium to GCP. The research validated that Radium achieved 30% higher Model FLOP Utilization (50-55% vs 39-45% for GCP) on identical JAX workloads.
- Massachusetts Institute of Technology (MIT)coreMIT researchers contributed to research suggesting that open models achieve roughly 80-90% of the performance of leading closed models at a fraction of the cost, which supports Radium's value proposition.
- Carnegie Mellon UniversitycoreCarnegie Mellon University conducted independent benchmarks comparing Radium to AWS. Research validated 2x lower inference latency vs AWS and 58% better utilization at 16-layer depth vs AWS, with the performance advantage compounding as models get deeper.
Scale indicators4 records
Recent moves6 records
Expansion highlights5 records
Radium competitors and assessment
Company assessmentDirect peers
- Fireworks AI: Fireworks AI provides serverless AI inference for open-source and fine-tuned models with OpenAI-compatible APIs, targeting enterprise developers seeking lower latency and cost. Directly comparable to Radium in pricing model, customer segment, and inference-optimization focus.
- OctoAI: OctoAI provides optimized AI inference infrastructure for running open and custom models at scale, with a focus on cost and latency reduction. Closely aligned with Radium's bare-metal optimization thesis and enterprise inference positioning.
- DeepInfra: DeepInfra offers low-cost, low-latency inference APIs for open-source LLMs and other models, targeting enterprise and developer use cases. Direct overlap with Radium on cost-optimized inference and OpenAI-compatible APIs.
- Together AI: Together AI operates an open-source-focused AI inference and training cloud with API endpoints positioned as lower-cost alternatives to OpenAI/Anthropic. It is the closest direct competitor to Radium in product, target customer (enterprise AI teams), and go-to-market motion (API-first, OpenAI-compatible).
- Modal: Modal provides serverless cloud infrastructure for AI/ML workloads with developer-focused APIs and auto-scaling. Comparable to Radium's serverless inference offering and emphasis on developer ergonomics for production AI deployments.
- Lambda: Lambda operates GPU cloud infrastructure purpose-built for AI training and inference, including reserved and on-demand GPU clusters. Comparable to Radium on the bare-metal AI infrastructure layer, though Lambda is more vertically integrated in hardware.
- Replicate: Replicate runs a cloud for running and deploying open-source machine learning models via API, with usage-based pricing similar to Radium's token-based model. Targets developers and enterprises seeking frictionless access to AI models without managing infrastructure.
- Anyscale: Anyscale offers AI compute and inference infrastructure built on Ray, serving enterprise teams running production AI workloads with a focus on performance and cost efficiency. Overlaps with Radium's enterprise AI infrastructure positioning and developer-first motion.
Broad incumbents
- CoreWeave: CoreWeave is a large-scale GPU cloud provider offering AI training and inference infrastructure to enterprise customers. It is a broader incumbent in the same category as Radium's underlying infrastructure, though CoreWeave's portfolio extends well beyond inference-specific workloads.
- AWS Bedrock: AWS Bedrock is a managed service providing access to multiple foundation models via API on AWS infrastructure. It is the broad incumbent competitor against which Radium benchmarks its 2x latency and 58% utilization advantages.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat5 records
Key risks5 records
Key highlights6 records
Customer concentration
Radium social profiles
Digital presenceRadium financial estimates
Financial estimateRevenue estimate
Valuation estimate
Radium leadership team
Management profileNumber of profiles
Profiles3 records
Radium funding detail
Funding detailFunding overview
Funding rounds4 records
Investors4 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Radium M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Radium
What does Radium do?
Radium is an OpenAI/Anthropic-compatible AI inference platform that delivers three open-source-based model tiers (Hal 1.0, Clarke 1.0, Tycho 1.0) over a serverless API. Customers access the models by swapping the endpoint URL—a single line of code—and pay per input and output token at approximately 50% of comparable frontier provider pricing, with optional prompt caching that reduces costs on repeated prompts by up to 90%.
Is Radium a public or private company?
Radium is a private company. It is classified as venture growth investor backed and is currently operating.
When was Radium founded?
Radium was founded in 2020. It employs 11 to 50 people.
Where is Radium based?
Radium is headquartered in San Francisco, United States, in the North America region.
How does Radium make money?
One revenue line is on record: API Token-based Inference.
Who are Radium's main competitors?
Direct peers on record are Fireworks AI, OctoAI, DeepInfra, Together AI, Modal, Lambda, Replicate and Anyscale. Broad incumbents are CoreWeave and AWS Bedrock.
Does Radium have an API?
Yes. Radium provides API-based access to AI model inference endpoints including Hal, Clarke, and Tycho model tiers. The API is OpenAI-compatible, allowing requests and responses to align with established integration patterns. Users can swap the endpoint URL with a single line of code to migrate from OpenAI or Anthropic. Supports token-based pricing per model tier with input and output token rates. Developer documentation is at radium.cloud/docs.