Developer docs
API playgroundTry for free, no card

Search company profiles

Cactus

Full company profile

uuid0002hbq

Namestring
Cactus
Legal namestring
Cactus Compute
Company typeenum
Private
Founded yearint
2025
Descriptiontext

Cactus Compute is a San Francisco-based, Y Combinator-backed developer infrastructure company founded in 2025 by Roman Shemet and Henry Ndubuaku. It builds an open-source, MIT-licensed on-device AI inference engine written in C/C++ that runs large language, speech, and vision models directly on consumer hardware, including iOS, Android, macOS, Linux, watchOS, and tvOS, with first-party SDKs in Swift, Kotlin, Flutter, React Native, Python, C++, and Rust. The runtime uses a custom CACT binary format, zero-copy mmap memory mapping, and hand-tuned ARM kernels (NEON, DOTPROD, I8MM, SME2), with native Apple Neural Engine support and Qualcomm NPU support planned; INT4/INT8 and the proprietary TurboQuant-H 2-bit quantization scheme are exposed to developers.

The product wraps a unified single-API surface around heterogeneous models — Whisper, Moonshine, NVIDIA Parakeet, Gemma 3/4, Qwen 3, and the LiquidAI LFM2 family — and adds a hybrid cloud-fallback router that offloads work to a managed backend when on-device capacity is exceeded. The company ships the in-house Needle 26M function-calling model and publishes research (TurboQuant-H) as part of its differentiation strategy.

Cactus monetizes through a freemium, product-led model: the engine and SDKs are free under MIT, while usage-based revenue is generated through the CACTUS_CLOUD_KEY metered cloud-handoff SKU and an enterprise intake path via Calendly. Distribution is community- and API-led, anchored by 4.2k+ GitHub stars and 53+ third-party comparison pages. The company has raised $500K from Y Combinator and FundersClub (Sept 2025) and operates with 1-10 employees, positioning it at the earliest commercialization stage.

Short descriptiontext

Cactus Compute builds an open-source, on-device AI inference engine with hybrid cloud fallback, serving mobile and edge developers across iOS, Android, macOS, Linux, watchOS, and tvOS via SDKs in seven languages. It monetizes usage-based cloud handoff under an MIT-licensed, product-led model.

Operating statusenum
Operating
Ownership categoryenum
Headcount rangeband
1–10
akta.pro rankint
HeadquartersSan Francisco, United States
HQ citystring
San Francisco
HQ countrystring
United States
HQ regionstring
North America
Markets served

Serves global market

Keyword5 values
on-device AI inference, edge AI runtime, mobile AI SDK, AI model quantization, hybrid cloud inference
Industry3 codes
1On-Device Inference Runtimes & SDKs (mobile/embedded)
CodeHDAAAJABPrimaryYes
2Model Deployment, Serving & Inference Platforms
CodeHDAAABAFPrimaryNo
3AI Compiler, Runtime & Kernel Optimization Software (CUDA/ROCm/XLA, graph compilers)
CodeHDAAAAAIPrimaryNo
NAICS code4 codes
  • Software Publishers5132
  • Software Publishers51321
  • Custom Computer Programming Services541511
  • Computer Systems Design and Related Services54151
SIC code3 codes
  • Services-Prepackaged Software7372
  • Services-Computer Programming Services7371
  • Services-Computer Integrated Systems Design7373
Product category
On-device AI Inference Engine
GTM motion3 records

Each record includes

Type, Description, Source

Revenue model2 records
1Cactus Hybrid Cloud API
TypeUsage Based
Description

Optional usage-based cloud API invoked automatically by the Cactus Hybrid Router when on-device confidence drops below a configured threshold. Provides a paid top-up on top of the free on-device engine.

cactuscompute.com
2Free On-Device Engine (MIT)
TypeFreemium
Description

The core Cactus inference engine and SDK are open-sourced under MIT and free to use, functioning as the freemium tier and adoption funnel; revenue is captured only when users enable optional cloud handoff or move to enterprise usage.

cactuscompute.com
Marketing channels7 records

Each record includes

Title, Type, Stage, Description, Source

Distribution channels5 records

Each record includes

Title, Type, Scope, Target buyer, Description, Source

Cost components5 values
Technology or R&D, Personnel, Infrastructure, Marketing or Sales, Operations
Pricing details1 tier
1Free on-device SDK + optional usage-based cloud handoff
ModelFreemiumBilling cadencePay-as-you-go
Notes

Core engine is MIT-licensed and free; cloud fallback requires CACT_CLOUD_KEY and is billed on usage. "Start for free" / "Free to start, scales with you" copy on the homepage.

cactuscompute.com
GTM typeB2B
B2B
Offering typeSoftware
Software
Brand1 of 3 records shown
1Cactus Hybrid Cloud
Description

Hybrid AI inference engine combining on-device inference with automatic cloud fallback for speech transcription, agents, and text generation.

cactuscompute.com
+2 more records
Core offering1 text field

Cactus Compute provides an open-source, MIT-licensed on-device AI inference engine (Cactus Engine) that loads, quantizes, and executes LLMs, speech-to-text, vision, and embedding models directly on smartphones, laptops, wearables, and edge hardware. A unified C FFI SDK exposes the engine across Python, Swift, Kotlin, Flutter, React Native, C++, and Rust for iOS, Android, macOS, Linux, watchOS, and tvOS. The optional Cactus Hybrid Cloud service adds confidence-based automatic cloud escalation for low-confidence requests, billed on usage.

Differentiator
Functional benefit
Problem solved
Product overview1 text field

Cactus Compute offers a unified on-device AI platform composed of a single core product, the Cactus Engine, accompanied by the optional Cactus Hybrid Cloud fallback service and the open-source Needle function-calling model. The Cactus Engine is the foundational inference runtime—an open-source, MIT-licensed engine that loads, quantizes, and executes LLMs, transcription, vision, and embedding models on smartphones, laptops, and edge hardware using its proprietary 'CACT' binary format. The cross-platform Cactus SDK (Swift, Kotlin, Flutter, React Native, Python, C++, Rust) and the Homebrew-distributed Cactus CLI expose the engine to developers, while Cactus Hybrid Cloud adds optional confidence-based cloud handoff for low-confidence requests. Needle, a 26M-parameter open-source function-calling model distilled from Gemini, ships as an example agentic model that demonstrates the engine's tool-use capability. Together these products form an end-to-end stack for running AI features on-device with automatic cloud escalation when needed.

Product and service5 records
1Cactus Engine
CategoryOn-device AI inference runtime
Description

An open-source, MIT-licensed on-device AI inference engine that loads, quantizes, and executes LLMs, transcription, vision, and embedding models on smartphones, laptops, and edge hardware. Uses zero-copy mmap loading with a proprietary CACT binary format and ships SIMD-optimized kernels for NEON, DOTPROD, I8MM, and SME2, targeting sub-120ms on-device latency and INT4/INT8 quantization.

2Cactus Hybrid Cloud
CategoryHybrid cloud AI inference API
Description

A confidence-based cloud fallback service that automatically routes low-confidence on-device requests to larger cloud models, providing cloud-level accuracy without always-on cloud cost. Accessed via the CACTUS_CLOUD_KEY environment variable and billed on usage. Marketed as delivering up to 5x cost savings versus cloud-only AI.

3Cactus SDK
CategoryDeveloper SDK / API library
Description

A cross-platform native SDK suite for integrating the Cactus Engine into mobile, desktop, and edge apps. Provides first-class bindings for Swift, Kotlin, Flutter, React Native, Python, C++, and Rust, built on top of a C FFI base layer, covering iOS, Android, macOS, Linux, watchOS, and tvOS from one SDK.

4Cactus CLI
CategoryDeveloper command-line tool
Description

A command-line interface for the Cactus Engine, installable via Homebrew on macOS, that supports downloading models from HuggingFace, converting them to the optimized CACT binary format, and starting interactive chat or transcription sessions.

5Needle
CategoryOpen-source AI model (function calling)
Description

An open-source 26M-parameter function-calling (tool use) model distilled from Gemini by Cactus Compute, using a 'Simple Attention Networks' architecture (attention and gating only, no MLPs). Pretrained on 200B tokens and post-trained on 2B tokens of synthesized function-calling data across 15 tool categories; runs at 6,000 tok/s prefill and 1,200 tok/s decode on consumer devices; MIT licensed.

Scale indicator10 records

Each record includes

Type, Value, Description, Source

Partnership8 partners
Strategic tierCoreTypeTechnology or Integration
Description

Cactus natively runs Liquid AI's LFM2 language models and LFM2-VL vision-language models on-device via the Cactus inference engine. The companies are explicitly described as complementary: Liquid AI builds efficient models, Cactus provides the mobile/edge runtime.

Strategic tierCoreTypeTechnology or Integration
Description

Cactus ships first-class support for NVIDIA's Parakeet CTC 1.1B non-autoregressive ASR model, with documented quickstart ('cactus transcribe nvidia/parakeet-ctc-1.1b'), Python/Rust/Swift/Kotlin/Flutter SDK examples, and benchmarked sub-200ms end-to-end latency on Apple Silicon.

Strategic tierCoreTypeTechnology or Integration
Description

Cactus supports Google's Gemma 3 / Gemma 4 E2B / Gemma 4 E4B / Gemma-270m models, including the AltUp per-layer embedding architectures used as the testbed for TurboQuant-H research. Models are downloadable directly via 'cactus download google/gemma-4-E2B-it'.

Strategic tierCoreTypeTechnology or Integration
Description

Cactus leverages Apple's Neural Engine for NPU acceleration on Apple devices and supports native Swift SDK + watchOS / tvOS deployment. Apple hardware (Vision Pro, iPhone 17 Pro, iPhone 13 Mini, M4 Pro) is a primary benchmarking and deployment target.

Strategic tierCoreTypeTechnology or Integration
Description

Cactus publishes open-source model weights on Hugging Face (e.g. Cactus-Compute/needle) and downloads models directly from the Hub via 'cactus download '.

Strategic tierCoreTypeTechnology or Integration
Description

The Cactus engine and Needle model are hosted on GitHub under the cactus-compute org, with 4.2k+ stars on the engine repo. GitHub is also used as the OAuth identity provider for cactuscompute.com signup and sign-in.

7AWS
Strategic tierMinorTypeOthers
Description

AWS is referenced only as an alumni origin of the founding team ('Built by a team from Y Combinator, University of Oxford, DeepRender, Salesforce, Google, AWS, Washington Post, MIT'); no commercial partnership or integration is described.

cactuscompute.com
Strategic tierMinorTypeOthers
Description

Google is referenced as a talent source for the founding team rather than as a commercial partner; a separate Google Gemma partnership entry covers model support.

Recent move6 records

Each record includes

Date, Type, Title, Description, Source

Expansion highlight5 records

Each record includes

Type, Description

Peers10 records
1llama.cpp
TypeDirect peer
Description

Open-source C/C++ LLM inference engine widely used for on-device and edge deployment. Directly comparable to the Cactus Engine, both targeting quantized LLM execution on consumer hardware with ARM SIMD optimization.

TypeDirect peer
Description

On-device AI platform offering an SDK and model hub for running LLMs, speech, and vision locally on phones, laptops, and edge devices. Direct competitor to Cactus in the on-device inference runtime category, explicitly listed in Cactus's comparison pages.

3ExecuTorch (Meta)
TypeDirect peer
Description

PyTorch's on-device inference runtime for mobile and edge, backed by Meta. Competes with Cactus for mobile LLM and multimodal deployment, especially on Android and embedded Linux.

4Core ML (Apple)
TypeBroad incumbent
Description

Apple's first-party on-device ML framework for iOS, macOS, watchOS, and tvOS. Comparable to Cactus's iOS/Swift story, but Apple-owned and bundled with the platform, making it the default incumbent Cactus differentiates against.

TypeBroad incumbent
Description

Google's widely deployed on-device inference runtime across Android and embedded. A broad incumbent Cactus competes with for mobile and edge AI workloads, and is explicitly featured in Cactus's comparison pages.

TypeBroad incumbent
Description

Cross-platform inference engine for ONNX models maintained by Microsoft and the community. Comparable broad incumbent for cross-platform on-device inference and one of the runtimes explicitly compared on Cactus's site.

7MLC LLM
TypeDirect peer
Description

Open-source project that compiles and deploys LLMs natively across iOS, Android, and WebGPU. Direct peer in the on-device LLM runtime category, also benchmarked and compared by Cactus.

8whisper.cpp
TypeDirect peer
Description

Open-source C/C++ port of OpenAI Whisper for on-device speech recognition, widely used in mobile transcription apps. Directly comparable to Cactus's transcription capability, and explicitly listed in Cactus's comparison pages.

TypeEmerging player
Description

Early-stage startup offering on-device speech recognition SDKs for mobile and edge. Direct emerging peer in the on-device ASR category that Cactus lists among its competitive comparisons.

TypeOthers
Description

Foundation-model company building efficient liquid neural networks (LFM2 series) that target on-device deployment. Strategic technology partner to Cactus, and an ecosystem participant whose model availability is core to Cactus's value proposition.

Market position
Strengths5 records

Each record includes

Headline, Details, Source

Weaknesses5 records

Each record includes

Headline, Details, Source

Competitive moat6 records

Each record includes

Type, Details

Key risks6 records

Each record includes

Headline, Details, Source

Key highlights7 records

Each record includes

Headline, Details, Source

Customer concentration

Classification, Details

Segment4 records

Each record includes

Title, Type, Primary, Description, Pain point addressed, Use case, Source

Ideal customer profile4 records

Each record includes

Profile, Firmographic size, Sales motion, Sales cycle length, Buying structure, Purchase trigger, Buyer persona, Geography, Industry vertical, Primary use case, Description, Pain points, Evidence proof points, Target buyer

Technology focused
Yes
API detail
Has APIbool
Yes

Docs URL, Description

Integration2 records

Each record includes

Title, Type, Description, Source

AI capability12 records

Each record includes

Type, Description, Source

AI maturity
App detail

Has app

Feature7 records

Each record includes

Title, Differentiator, Description, Source

Core technology
Revenue estimate
Valuation estimate
Number of profiles
Profiles2 records

Each record includes

Name, Designation, Designation category, Overview, Profile commentary, Source

No data
Compliance2 records

Each record includes

Name, Class, Description

Funding overview

Funding stage, Last funding date, Total funding USD

Funding rounds2 records

Each record includes

Round, Amount USD, Date, Pre money valuation, Total investors, Investors, News

Investors3 records

Each record includes

Name, Type, Date of entry, Rounds participated, Website

Funding detail is available on the Subscription and Enterprise plan.Contact sales →

M&A

Each record includes

Name, Acquisition type, Announced date, Completed date, Status, Website, News

Investment

Each record includes

Name, Round, Announced date, Lead investor, Website, News

M&A and investment is available on the Subscription and Enterprise plan.Contact sales →

Cactus

On-device AI Inference Enginecactuscompute.com

Cactus Compute builds an open-source, on-device AI inference engine with hybrid cloud fallback, serving mobile and edge developers across iOS, Android, macOS, Linux, watchOS, and tvOS via SDKs in seven languages. It monetizes usage-based cloud handoff under an MIT-licensed, product-led model.

What Cactus does

Cactus Compute is a San Francisco-based, Y Combinator-backed developer infrastructure company founded in 2025 by Roman Shemet and Henry Ndubuaku. It builds an open-source, MIT-licensed on-device AI inference engine written in C/C++ that runs large language, speech, and vision models directly on consumer hardware, including iOS, Android, macOS, Linux, watchOS, and tvOS, with first-party SDKs in Swift, Kotlin, Flutter, React Native, Python, C++, and Rust. The runtime uses a custom CACT binary format, zero-copy mmap memory mapping, and hand-tuned ARM kernels (NEON, DOTPROD, I8MM, SME2), with native Apple Neural Engine support and Qualcomm NPU support planned; INT4/INT8 and the proprietary TurboQuant-H 2-bit quantization scheme are exposed to developers.

The product wraps a unified single-API surface around heterogeneous models — Whisper, Moonshine, NVIDIA Parakeet, Gemma 3/4, Qwen 3, and the LiquidAI LFM2 family — and adds a hybrid cloud-fallback router that offloads work to a managed backend when on-device capacity is exceeded. The company ships the in-house Needle 26M function-calling model and publishes research (TurboQuant-H) as part of its differentiation strategy.

Cactus monetizes through a freemium, product-led model: the engine and SDKs are free under MIT, while usage-based revenue is generated through the CACTUS_CLOUD_KEY metered cloud-handoff SKU and an enterprise intake path via Calendly. Distribution is community- and API-led, anchored by 4.2k+ GitHub stars and 53+ third-party comparison pages. The company has raised $500K from Y Combinator and FundersClub (Sept 2025) and operates with 1-10 employees, positioning it at the earliest commercialization stage.

Cactus firmographics

Firmographics
Name
Cactus
Legal name
Cactus Compute
Website
https://cactuscompute.com
Company type
Private
Founded year
2025
Operating status
Operating
Headcount range
1–10 employees
Short description
Cactus Compute builds an open-source, on-device AI inference engine with hybrid cloud fallback, serving mobile and edge developers across iOS, Android, macOS, Linux, watchOS, and tvOS via SDKs in seven languages. It monetizes usage-based cloud handoff under an MIT-licensed, product-led model.
Ownership category
akta.pro rank

Cactus industry classification

Industry
Product category
On-device AI Inference Engine
NAICS
Software Publishers (5132), Software Publishers (51321), Custom Computer Programming Services (541511), Computer Systems Design and Related Services (54151)
SIC
Services-Prepackaged Software (7372), Services-Computer Programming Services (7371), Services-Computer Integrated Systems Design (7373)
akta.pro primary industry
On-Device Inference Runtimes & SDKs (mobile/embedded) (HDAAAJAB)
akta.pro secondary industries
Model Deployment, Serving & Inference Platforms (HDAAABAF), AI Compiler, Runtime & Kernel Optimization Software (CUDA/ROCm/XLA, graph compilers) (HDAAAAAI)

Keywords

  • On-device AI inference
  • Edge AI runtime
  • Mobile AI SDK
  • AI model quantization
  • Hybrid cloud inference

Where Cactus is headquartered

Location

Headquarters

HQ city
San Francisco
HQ country
United States
HQ region
North America

Markets served

Cactus business model

Business model
GTM type
B2B
Offering type
Software
Cost components
Technology or R&D, Personnel, Infrastructure, Marketing or Sales, Operations

Revenue model

  1. Cactus Hybrid Cloud API: Optional usage-based cloud API invoked automatically by the Cactus Hybrid Router when on-device confidence drops below a configured threshold. Provides a paid top-up on top of the free on-device engine.
  2. Free On-Device Engine (MIT): The core Cactus inference engine and SDK are open-sourced under MIT and free to use, functioning as the freemium tier and adoption funnel; revenue is captured only when users enable optional cloud handoff or move to enterprise usage.

Pricing tiers

ModelBillingPrice
FreemiumPay-as-you-goFree on-device SDK + optional usage-based cloud handoff

Go-to-market motion3 records

Distribution channels5 records

Marketing channels7 records

Cactus product offering

Product offering

Core offering

Cactus Compute provides an open-source, MIT-licensed on-device AI inference engine (Cactus Engine) that loads, quantizes, and executes LLMs, speech-to-text, vision, and embedding models directly on smartphones, laptops, wearables, and edge hardware. A unified C FFI SDK exposes the engine across Python, Swift, Kotlin, Flutter, React Native, C++, and Rust for iOS, Android, macOS, Linux, watchOS, and tvOS. The optional Cactus Hybrid Cloud service adds confidence-based automatic cloud escalation for low-confidence requests, billed on usage.

Product overview

Cactus Compute offers a unified on-device AI platform composed of a single core product, the Cactus Engine, accompanied by the optional Cactus Hybrid Cloud fallback service and the open-source Needle function-calling model. The Cactus Engine is the foundational inference runtime—an open-source, MIT-licensed engine that loads, quantizes, and executes LLMs, transcription, vision, and embedding models on smartphones, laptops, and edge hardware using its proprietary 'CACT' binary format. The cross-platform Cactus SDK (Swift, Kotlin, Flutter, React Native, Python, C++, Rust) and the Homebrew-distributed Cactus CLI expose the engine to developers, while Cactus Hybrid Cloud adds optional confidence-based cloud handoff for low-confidence requests. Needle, a 26M-parameter open-source function-calling model distilled from Gemini, ships as an example agentic model that demonstrates the engine's tool-use capability. Together these products form an end-to-end stack for running AI features on-device with automatic cloud escalation when needed.

Differentiator

Problem solved

Functional benefit

Brands

  • Cactus Hybrid Cloud: Hybrid AI inference engine combining on-device inference with automatic cloud fallback for speech transcription, agents, and text generation.
  • Cactus Engine
  • Needle

Products and services

  • Cactus Engine An open-source, MIT-licensed on-device AI inference engine that loads, quantizes, and executes LLMs, transcription, vision, and embedding models on smartphones, laptops, and edge hardware. Uses zero-copy mmap loading with a proprietary CACT binary format and ships SIMD-optimized kernels for NEON, DOTPROD, I8MM, and SME2, targeting sub-120ms on-device latency and INT4/INT8 quantization.
  • Cactus Hybrid Cloud A confidence-based cloud fallback service that automatically routes low-confidence on-device requests to larger cloud models, providing cloud-level accuracy without always-on cloud cost. Accessed via the CACTUS_CLOUD_KEY environment variable and billed on usage. Marketed as delivering up to 5x cost savings versus cloud-only AI.
  • Cactus SDK A cross-platform native SDK suite for integrating the Cactus Engine into mobile, desktop, and edge apps. Provides first-class bindings for Swift, Kotlin, Flutter, React Native, Python, C++, and Rust, built on top of a C FFI base layer, covering iOS, Android, macOS, Linux, watchOS, and tvOS from one SDK.
  • Cactus CLI A command-line interface for the Cactus Engine, installable via Homebrew on macOS, that supports downloading models from HuggingFace, converting them to the optimized CACT binary format, and starting interactive chat or transcription sessions.
  • Needle An open-source 26M-parameter function-calling (tool use) model distilled from Gemini by Cactus Compute, using a 'Simple Attention Networks' architecture (attention and gating only, no MLPs). Pretrained on 200B tokens and post-trained on 2B tokens of synthesized function-calling data across 15 tool categories; runs at 6,000 tok/s prefill and 1,200 tok/s decode on consumer devices; MIT licensed.

Companies that use Cactus

Customer profile

Segments4 records

Ideal customer profiles4 records

Cactus technology and API

Technology

Technology focussed Yes

API detail

Has API
Yes
API docs
API detail

Core technology

AI maturity

App detail

Integration2 records

AI capability12 records

Feature7 records

Cactus partnerships and signals

Strategic signal

Partnerships

Eight partnerships are on record, tiered core and minor.

  • Liquid AIcoreTechnology or IntegrationCactus natively runs Liquid AI's LFM2 language models and LFM2-VL vision-language models on-device via the Cactus inference engine. The companies are explicitly described as complementary: Liquid AI builds efficient models, Cactus provides the mobile/edge runtime.
  • NVIDIAcoreTechnology or IntegrationCactus ships first-class support for NVIDIA's Parakeet CTC 1.1B non-autoregressive ASR model, with documented quickstart ('cactus transcribe nvidia/parakeet-ctc-1.1b'), Python/Rust/Swift/Kotlin/Flutter SDK examples, and benchmarked sub-200ms end-to-end latency on Apple Silicon.
  • Google (Gemma)coreTechnology or IntegrationCactus supports Google's Gemma 3 / Gemma 4 E2B / Gemma 4 E4B / Gemma-270m models, including the AltUp per-layer embedding architectures used as the testbed for TurboQuant-H research. Models are downloadable directly via 'cactus download google/gemma-4-E2B-it'.
  • ApplecoreTechnology or IntegrationCactus leverages Apple's Neural Engine for NPU acceleration on Apple devices and supports native Swift SDK + watchOS / tvOS deployment. Apple hardware (Vision Pro, iPhone 17 Pro, iPhone 13 Mini, M4 Pro) is a primary benchmarking and deployment target.
  • Hugging FacecoreTechnology or IntegrationCactus publishes open-source model weights on Hugging Face (e.g. Cactus-Compute/needle) and downloads models directly from the Hub via 'cactus download '.
  • GitHubcoreTechnology or IntegrationThe Cactus engine and Needle model are hosted on GitHub under the cactus-compute org, with 4.2k+ stars on the engine repo. GitHub is also used as the OAuth identity provider for cactuscompute.com signup and sign-in.
  • AWSminorOthersAWS is referenced only as an alumni origin of the founding team ('Built by a team from Y Combinator, University of Oxford, DeepRender, Salesforce, Google, AWS, Washington Post, MIT'); no commercial partnership or integration is described.
  • Google (alumni)minorOthersGoogle is referenced as a talent source for the founding team rather than as a commercial partner; a separate Google Gemma partnership entry covers model support.

Scale indicators10 records

Recent moves6 records

Expansion highlights5 records

Cactus competitors and assessment

Company assessment

Direct peers

  • llama.cpp: Open-source C/C++ LLM inference engine widely used for on-device and edge deployment. Directly comparable to the Cactus Engine, both targeting quantized LLM execution on consumer hardware with ARM SIMD optimization.
  • Nexa AI: On-device AI platform offering an SDK and model hub for running LLMs, speech, and vision locally on phones, laptops, and edge devices. Direct competitor to Cactus in the on-device inference runtime category, explicitly listed in Cactus's comparison pages.
  • ExecuTorch (Meta): PyTorch's on-device inference runtime for mobile and edge, backed by Meta. Competes with Cactus for mobile LLM and multimodal deployment, especially on Android and embedded Linux.
  • MLC LLM: Open-source project that compiles and deploys LLMs natively across iOS, Android, and WebGPU. Direct peer in the on-device LLM runtime category, also benchmarked and compared by Cactus.
  • whisper.cpp: Open-source C/C++ port of OpenAI Whisper for on-device speech recognition, widely used in mobile transcription apps. Directly comparable to Cactus's transcription capability, and explicitly listed in Cactus's comparison pages.

Broad incumbents

  • Core ML (Apple): Apple's first-party on-device ML framework for iOS, macOS, watchOS, and tvOS. Comparable to Cactus's iOS/Swift story, but Apple-owned and bundled with the platform, making it the default incumbent Cactus differentiates against.
  • TensorFlow Lite: Google's widely deployed on-device inference runtime across Android and embedded. A broad incumbent Cactus competes with for mobile and edge AI workloads, and is explicitly featured in Cactus's comparison pages.
  • ONNX Runtime: Cross-platform inference engine for ONNX models maintained by Microsoft and the community. Comparable broad incumbent for cross-platform on-device inference and one of the runtimes explicitly compared on Cactus's site.

Emerging players

  • Argmax: Early-stage startup offering on-device speech recognition SDKs for mobile and edge. Direct emerging peer in the on-device ASR category that Cactus lists among its competitive comparisons.

Others

  • Liquid AI: Foundation-model company building efficient liquid neural networks (LFM2 series) that target on-device deployment. Strategic technology partner to Cactus, and an ecosystem participant whose model availability is core to Cactus's value proposition.

Market position

Strengths5 records

Weaknesses5 records

Competitive moat6 records

Key risks6 records

Key highlights7 records

Customer concentration

Cactus social profiles

Digital presence

Cactus compliance and trust

Trust signal

Compliance2 records

Cactus financial estimates

Financial estimate

Revenue estimate

Valuation estimate

Cactus leadership team

Management profile

Number of profiles

Profiles2 records

Cactus funding detail

Funding detail

Funding overview

Funding rounds2 records

Investors3 records

Funding detail is available on the Subscription and Enterprise plan.Contact sales →

Cactus M&A and investment

M&A and investment

M&A

Investments

M&A and investment is available on the Subscription and Enterprise plan.Contact sales →

Frequently asked questions about Cactus

What does Cactus do?

Cactus Compute provides an open-source, MIT-licensed on-device AI inference engine (Cactus Engine) that loads, quantizes, and executes LLMs, speech-to-text, vision, and embedding models directly on smartphones, laptops, wearables, and edge hardware. A unified C FFI SDK exposes the engine across Python, Swift, Kotlin, Flutter, React Native, C++, and Rust for iOS, Android, macOS, Linux, watchOS, and tvOS. The optional Cactus Hybrid Cloud service adds confidence-based automatic cloud escalation for low-confidence requests, billed on usage.

Is Cactus a public or private company?

Cactus is a private company. It is classified as venture growth investor backed and is currently operating.

When was Cactus founded?

Cactus was founded in 2025. It employs 1 to 10 people.

Where is Cactus based?

Cactus is headquartered in San Francisco, United States, in the North America region.

How does Cactus make money?

Two revenue lines are on record. Cactus Hybrid Cloud API is the primary driver. The others are free On-Device Engine (MIT).

Who are Cactus's main competitors?

Direct peers on record are llama.cpp, Nexa AI, ExecuTorch (Meta), MLC LLM and whisper.cpp. Broad incumbents are Core ML (Apple), TensorFlow Lite and ONNX Runtime. Argmax is listed as an emerging player. Liquid AI is listed as an others.

Does Cactus have an API?

Yes. Cactus exposes a C FFI base layer (libcactus) with native high-level SDK bindings for Python, Swift, Kotlin, Flutter, React Native, C++, and Rust. Developers can initialize models, run chat completions with streaming/token callbacks, perform file-based and real-time microphone transcription, and invoke function/tool calling. The Cactus Hybrid Cloud API is accessed via an API key supplied through the CACTUS_CLOUD_KEY environment variable and is used to route low-confidence on-device requests to the cloud for improved accuracy. Responses can carry a cloud_handoff flag and confidence_threshold is configurable. Cactus provides streaming transcription, async audio callbacks, and unified APIs across LLM, transcription, vision, and embeddings modalities. Developer documentation is at docs.cactuscompute.com.

What industry is Cactus in?

Cactus's product category is On-device AI Inference Engine. Its primary akta.pro industry code is HDAAAJAB, On-Device Inference Runtimes & SDKs (mobile/embedded), with a secondary code of HDAAABAF, Model Deployment, Serving & Inference Platforms. Its NAICS code is 5132 and its SIC code is 7372.

Unlock the full company data

50 free credits on sign-up, no credit card required.

Contact sales
Live signals
Geeky GadgetsNeedle 2 Local AI Packs a 45-Million Parameter LLM Into Just 14MBCactus Compute has developed Needle 2, a compact 14MB agentic large language model with 45 million parameters designed to run on resource-constrained hardware like the ESP32S3 microcontroller. The model specializes in converting natural language prompts into precise device actions for applications such as smart home automation and robotics control, achieving high accuracy through techniques like two-bit quantization. Its offline functionality and open-source Apache 2.0 license position it as a specialized tool for edge computing environments where efficiency and security are paramount.MarkTechPostMeet Needle 2: An Open 45M-Parameter Tool-Calling Model That Ships as a 14MB Binary and Runs a Full Session in 28MB of RAMCactus Compute has released Needle 2, an open 45M-parameter model designed for tool calling and application on constrained hardware environments. The model operates within a compact 14MB binary and requires approximately 28MB of RAM to run a full session, achieving notable decode throughput rates on various devices. Needle 2 is positioned to serve industries needing offline processing for device interactions, such as smart home and wearable technologies.American Banking and Market News9,160 Shares in Cactus, Inc. $WHD Acquired by AmundiAmundi acquired 9,160 shares of Cactus, Inc. valued at approximately $434,000 during the first quarter, according to a recent SEC filing. Cactus also recently reported strong quarterly earnings, with a revenue increase of 64.3% year-over-year and announced a dividend increase from $0.14 to $0.15 per share.AInvestIs Cactus Worth Buying as Growth Accelerates but Valuation Stretches?Cactus, Inc. is showing accelerated earnings growth with international expansion and substantial liquidity, driven by acquisitions and increased international orders. The company faces valuation premiums, tariff-related risks, and margin pressures, which require steady earnings to justify its valuation. Its outlook remains positive with focused growth despite operational risks.Startup FortuneNeedle shows tiny models can move AI agents onto devicesCactus Compute has open-sourced Needle, a 26-million-parameter tool-calling model designed to run locally on consumer devices like phones and watches. The model utilizes a unique Simple Attention Network architecture and was trained using synthetic data generated by Google's Gemini to enable efficient, low-latency function calling without relying on cloud inference.MediaPostWright-Patt Credit Union Banks On Cactus As AORCactus was selected by Wright-Patt Credit Union as integrated AOR to grow its member base and expand into new markets. First campaign work is expected in June. The agency's clients include Fjällräven and Colorado Lottery.Business Wire BlogCactus Announces Board and Executive Leadership TransitionsCactus announced Tana Utley's election to its Board of Directors at the May 12, 2026 annual meeting, reducing the board to eight members. Steven Bender was appointed CEO of the Spoolable Technologies Segment, taking over from Stephen Tadlock. The company also noted Bruce Rothstein and Melissa Law declined reelection.Business Wire BlogCactus Announces First Quarter 2026 ResultsCactus reported Q1 2026 revenue of $388.3 million and operating income of $49.5 million, with net income of $40.2 million. The company closed its acquisition of a majority interest in Baker Hughes' Surface Pressure Control business on January 1, 2026, and expects Q2 Pressure Control revenues to be flat amid Middle East conflict impacts.BitgetTop Position at 10%: What Makes This Almost $60 Million Investment in an Oilfield Company RemarkableWebs Creek Capital Management disclosed a new investment in Cactus (NYSE:WHD), acquiring 1,263,873 shares valued at approximately $57.73 million during the fourth quarter, representing 10.33% of the firm's assets under management and making it the largest holding in its portfolio. Cactus, which designs, produces, and leases wellheads and pressure control equipment for unconventional oil and gas wells, reported quarterly revenue of $261 million with an 18.5% net income margin, though annual revenue declined from $1.13 billion to $1.08 billion. The article notes that Cactus occupies a distinct segment of the energy value chain insulated from direct oil price fluctuations but faces margin compression and slowing growth, with the pending acquisition of Baker Hughes' surface pressure control business positioned as a potential growth catalyst.PR NewswireCactus to Open Restaurant at The Harvest YardCactus, a Mexican restaurant brand with six existing locations in the Puget Sound area, has announced plans to open a new restaurant at The Harvest Yard, a mixed-use development in Woodinville, Washington, scheduled for summer 2026. The Harvest Yard development also includes The SOMM Hotel & Spa (part of Marriott's Autograph Collection), multiple restaurants, tasting rooms, retail, and residential units. Other committed tenants include Crawlspace Gastropub and Dossier Wine Collective, among others, positioning the development as a regional food and wine destination.