Cactus
Cactus Compute builds an open-source, on-device AI inference engine with hybrid cloud fallback, serving mobile and edge developers across iOS, Android, macOS, Linux, watchOS, and tvOS via SDKs in seven languages. It monetizes usage-based cloud handoff under an MIT-licensed, product-led model.
- Company typePrivate
- Founded2025
- HeadquartersSan Francisco, United States
- Headcount1–10
- GTM typeB2B
- OfferingSoftware
What Cactus does
Cactus Compute is a San Francisco-based, Y Combinator-backed developer infrastructure company founded in 2025 by Roman Shemet and Henry Ndubuaku. It builds an open-source, MIT-licensed on-device AI inference engine written in C/C++ that runs large language, speech, and vision models directly on consumer hardware, including iOS, Android, macOS, Linux, watchOS, and tvOS, with first-party SDKs in Swift, Kotlin, Flutter, React Native, Python, C++, and Rust. The runtime uses a custom CACT binary format, zero-copy mmap memory mapping, and hand-tuned ARM kernels (NEON, DOTPROD, I8MM, SME2), with native Apple Neural Engine support and Qualcomm NPU support planned; INT4/INT8 and the proprietary TurboQuant-H 2-bit quantization scheme are exposed to developers.
The product wraps a unified single-API surface around heterogeneous models — Whisper, Moonshine, NVIDIA Parakeet, Gemma 3/4, Qwen 3, and the LiquidAI LFM2 family — and adds a hybrid cloud-fallback router that offloads work to a managed backend when on-device capacity is exceeded. The company ships the in-house Needle 26M function-calling model and publishes research (TurboQuant-H) as part of its differentiation strategy.
Cactus monetizes through a freemium, product-led model: the engine and SDKs are free under MIT, while usage-based revenue is generated through the CACTUS_CLOUD_KEY metered cloud-handoff SKU and an enterprise intake path via Calendly. Distribution is community- and API-led, anchored by 4.2k+ GitHub stars and 53+ third-party comparison pages. The company has raised $500K from Y Combinator and FundersClub (Sept 2025) and operates with 1-10 employees, positioning it at the earliest commercialization stage.
Cactus firmographics
Firmographics- Name
- Cactus
- Legal name
- Cactus Compute
- Website
- https://cactuscompute.com
- Company type
- Private
- Founded year
- 2025
- Operating status
- Operating
- Headcount range
- 1–10 employees
- Short description
- Cactus Compute builds an open-source, on-device AI inference engine with hybrid cloud fallback, serving mobile and edge developers across iOS, Android, macOS, Linux, watchOS, and tvOS via SDKs in seven languages. It monetizes usage-based cloud handoff under an MIT-licensed, product-led model.
- Ownership category
- akta.pro rank
Cactus industry classification
Industry- Product category
- On-device AI Inference Engine
- NAICS
- Software Publishers (5132), Software Publishers (51321), Custom Computer Programming Services (541511), Computer Systems Design and Related Services (54151)
- SIC
- Services-Prepackaged Software (7372), Services-Computer Programming Services (7371), Services-Computer Integrated Systems Design (7373)
- akta.pro primary industry
- On-Device Inference Runtimes & SDKs (mobile/embedded) (HDAAAJAB)
- akta.pro secondary industries
- Model Deployment, Serving & Inference Platforms (HDAAABAF), AI Compiler, Runtime & Kernel Optimization Software (CUDA/ROCm/XLA, graph compilers) (HDAAAAAI)
Keywords
Where Cactus is headquartered
LocationHeadquarters
- HQ city
- San Francisco
- HQ country
- United States
- HQ region
- North America
Markets served
Cactus business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Technology or R&D, Personnel, Infrastructure, Marketing or Sales, Operations
Revenue model
- Cactus Hybrid Cloud API: Optional usage-based cloud API invoked automatically by the Cactus Hybrid Router when on-device confidence drops below a configured threshold. Provides a paid top-up on top of the free on-device engine.
- Free On-Device Engine (MIT): The core Cactus inference engine and SDK are open-sourced under MIT and free to use, functioning as the freemium tier and adoption funnel; revenue is captured only when users enable optional cloud handoff or move to enterprise usage.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Freemium | Pay-as-you-go | Free on-device SDK + optional usage-based cloud handoff |
Go-to-market motion3 records
Distribution channels5 records
Marketing channels7 records
Cactus product offering
Product offeringCore offering
Cactus Compute provides an open-source, MIT-licensed on-device AI inference engine (Cactus Engine) that loads, quantizes, and executes LLMs, speech-to-text, vision, and embedding models directly on smartphones, laptops, wearables, and edge hardware. A unified C FFI SDK exposes the engine across Python, Swift, Kotlin, Flutter, React Native, C++, and Rust for iOS, Android, macOS, Linux, watchOS, and tvOS. The optional Cactus Hybrid Cloud service adds confidence-based automatic cloud escalation for low-confidence requests, billed on usage.
Product overview
Cactus Compute offers a unified on-device AI platform composed of a single core product, the Cactus Engine, accompanied by the optional Cactus Hybrid Cloud fallback service and the open-source Needle function-calling model. The Cactus Engine is the foundational inference runtime—an open-source, MIT-licensed engine that loads, quantizes, and executes LLMs, transcription, vision, and embedding models on smartphones, laptops, and edge hardware using its proprietary 'CACT' binary format. The cross-platform Cactus SDK (Swift, Kotlin, Flutter, React Native, Python, C++, Rust) and the Homebrew-distributed Cactus CLI expose the engine to developers, while Cactus Hybrid Cloud adds optional confidence-based cloud handoff for low-confidence requests. Needle, a 26M-parameter open-source function-calling model distilled from Gemini, ships as an example agentic model that demonstrates the engine's tool-use capability. Together these products form an end-to-end stack for running AI features on-device with automatic cloud escalation when needed.
Differentiator
Problem solved
Functional benefit
Brands
- Cactus Hybrid Cloud: Hybrid AI inference engine combining on-device inference with automatic cloud fallback for speech transcription, agents, and text generation.
- Cactus Engine
- Needle
Products and services
- Cactus Engine An open-source, MIT-licensed on-device AI inference engine that loads, quantizes, and executes LLMs, transcription, vision, and embedding models on smartphones, laptops, and edge hardware. Uses zero-copy mmap loading with a proprietary CACT binary format and ships SIMD-optimized kernels for NEON, DOTPROD, I8MM, and SME2, targeting sub-120ms on-device latency and INT4/INT8 quantization.
- Cactus Hybrid Cloud A confidence-based cloud fallback service that automatically routes low-confidence on-device requests to larger cloud models, providing cloud-level accuracy without always-on cloud cost. Accessed via the CACTUS_CLOUD_KEY environment variable and billed on usage. Marketed as delivering up to 5x cost savings versus cloud-only AI.
- Cactus SDK A cross-platform native SDK suite for integrating the Cactus Engine into mobile, desktop, and edge apps. Provides first-class bindings for Swift, Kotlin, Flutter, React Native, Python, C++, and Rust, built on top of a C FFI base layer, covering iOS, Android, macOS, Linux, watchOS, and tvOS from one SDK.
- Cactus CLI A command-line interface for the Cactus Engine, installable via Homebrew on macOS, that supports downloading models from HuggingFace, converting them to the optimized CACT binary format, and starting interactive chat or transcription sessions.
- Needle An open-source 26M-parameter function-calling (tool use) model distilled from Gemini by Cactus Compute, using a 'Simple Attention Networks' architecture (attention and gating only, no MLPs). Pretrained on 200B tokens and post-trained on 2B tokens of synthesized function-calling data across 15 tool categories; runs at 6,000 tok/s prefill and 1,200 tok/s decode on consumer devices; MIT licensed.
Companies that use Cactus
Customer profileSegments4 records
Ideal customer profiles4 records
Cactus technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration2 records
AI capability12 records
Feature7 records
Cactus partnerships and signals
Strategic signalPartnerships
Eight partnerships are on record, tiered core and minor.
- Liquid AIcoreCactus natively runs Liquid AI's LFM2 language models and LFM2-VL vision-language models on-device via the Cactus inference engine. The companies are explicitly described as complementary: Liquid AI builds efficient models, Cactus provides the mobile/edge runtime.
- NVIDIAcoreCactus ships first-class support for NVIDIA's Parakeet CTC 1.1B non-autoregressive ASR model, with documented quickstart ('cactus transcribe nvidia/parakeet-ctc-1.1b'), Python/Rust/Swift/Kotlin/Flutter SDK examples, and benchmarked sub-200ms end-to-end latency on Apple Silicon.
- Google (Gemma)coreCactus supports Google's Gemma 3 / Gemma 4 E2B / Gemma 4 E4B / Gemma-270m models, including the AltUp per-layer embedding architectures used as the testbed for TurboQuant-H research. Models are downloadable directly via 'cactus download google/gemma-4-E2B-it'.
- ApplecoreCactus leverages Apple's Neural Engine for NPU acceleration on Apple devices and supports native Swift SDK + watchOS / tvOS deployment. Apple hardware (Vision Pro, iPhone 17 Pro, iPhone 13 Mini, M4 Pro) is a primary benchmarking and deployment target.
- Hugging FacecoreCactus publishes open-source model weights on Hugging Face (e.g. Cactus-Compute/needle) and downloads models directly from the Hub via 'cactus download '.
- GitHubcoreThe Cactus engine and Needle model are hosted on GitHub under the cactus-compute org, with 4.2k+ stars on the engine repo. GitHub is also used as the OAuth identity provider for cactuscompute.com signup and sign-in.
- AWSminorAWS is referenced only as an alumni origin of the founding team ('Built by a team from Y Combinator, University of Oxford, DeepRender, Salesforce, Google, AWS, Washington Post, MIT'); no commercial partnership or integration is described.
- Google (alumni)minorGoogle is referenced as a talent source for the founding team rather than as a commercial partner; a separate Google Gemma partnership entry covers model support.
Scale indicators10 records
Recent moves6 records
Expansion highlights5 records
Cactus competitors and assessment
Company assessmentDirect peers
- llama.cpp: Open-source C/C++ LLM inference engine widely used for on-device and edge deployment. Directly comparable to the Cactus Engine, both targeting quantized LLM execution on consumer hardware with ARM SIMD optimization.
- Nexa AI: On-device AI platform offering an SDK and model hub for running LLMs, speech, and vision locally on phones, laptops, and edge devices. Direct competitor to Cactus in the on-device inference runtime category, explicitly listed in Cactus's comparison pages.
- ExecuTorch (Meta): PyTorch's on-device inference runtime for mobile and edge, backed by Meta. Competes with Cactus for mobile LLM and multimodal deployment, especially on Android and embedded Linux.
- MLC LLM: Open-source project that compiles and deploys LLMs natively across iOS, Android, and WebGPU. Direct peer in the on-device LLM runtime category, also benchmarked and compared by Cactus.
- whisper.cpp: Open-source C/C++ port of OpenAI Whisper for on-device speech recognition, widely used in mobile transcription apps. Directly comparable to Cactus's transcription capability, and explicitly listed in Cactus's comparison pages.
Broad incumbents
- Core ML (Apple): Apple's first-party on-device ML framework for iOS, macOS, watchOS, and tvOS. Comparable to Cactus's iOS/Swift story, but Apple-owned and bundled with the platform, making it the default incumbent Cactus differentiates against.
- TensorFlow Lite: Google's widely deployed on-device inference runtime across Android and embedded. A broad incumbent Cactus competes with for mobile and edge AI workloads, and is explicitly featured in Cactus's comparison pages.
- ONNX Runtime: Cross-platform inference engine for ONNX models maintained by Microsoft and the community. Comparable broad incumbent for cross-platform on-device inference and one of the runtimes explicitly compared on Cactus's site.
Emerging players
- Argmax: Early-stage startup offering on-device speech recognition SDKs for mobile and edge. Direct emerging peer in the on-device ASR category that Cactus lists among its competitive comparisons.
Others
- Liquid AI: Foundation-model company building efficient liquid neural networks (LFM2 series) that target on-device deployment. Strategic technology partner to Cactus, and an ecosystem participant whose model availability is core to Cactus's value proposition.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat6 records
Key risks6 records
Key highlights7 records
Customer concentration
Cactus social profiles
Digital presenceCactus compliance and trust
Trust signalCompliance2 records
Cactus financial estimates
Financial estimateRevenue estimate
Valuation estimate
Cactus leadership team
Management profileNumber of profiles
Profiles2 records
Cactus funding detail
Funding detailFunding overview
Funding rounds2 records
Investors3 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Cactus M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Cactus
What does Cactus do?
Cactus Compute provides an open-source, MIT-licensed on-device AI inference engine (Cactus Engine) that loads, quantizes, and executes LLMs, speech-to-text, vision, and embedding models directly on smartphones, laptops, wearables, and edge hardware. A unified C FFI SDK exposes the engine across Python, Swift, Kotlin, Flutter, React Native, C++, and Rust for iOS, Android, macOS, Linux, watchOS, and tvOS. The optional Cactus Hybrid Cloud service adds confidence-based automatic cloud escalation for low-confidence requests, billed on usage.
Is Cactus a public or private company?
Cactus is a private company. It is classified as venture growth investor backed and is currently operating.
When was Cactus founded?
Cactus was founded in 2025. It employs 1 to 10 people.
Where is Cactus based?
Cactus is headquartered in San Francisco, United States, in the North America region.
How does Cactus make money?
Two revenue lines are on record. Cactus Hybrid Cloud API is the primary driver. The others are free On-Device Engine (MIT).
Who are Cactus's main competitors?
Direct peers on record are llama.cpp, Nexa AI, ExecuTorch (Meta), MLC LLM and whisper.cpp. Broad incumbents are Core ML (Apple), TensorFlow Lite and ONNX Runtime. Argmax is listed as an emerging player. Liquid AI is listed as an others.
Does Cactus have an API?
Yes. Cactus exposes a C FFI base layer (libcactus) with native high-level SDK bindings for Python, Swift, Kotlin, Flutter, React Native, C++, and Rust. Developers can initialize models, run chat completions with streaming/token callbacks, perform file-based and real-time microphone transcription, and invoke function/tool calling. The Cactus Hybrid Cloud API is accessed via an API key supplied through the CACTUS_CLOUD_KEY environment variable and is used to route low-confidence on-device requests to the cloud for improved accuracy. Responses can carry a cloud_handoff flag and confidence_threshold is configurable. Cactus provides streaming transcription, async audio callbacks, and unified APIs across LLM, transcription, vision, and embeddings modalities. Developer documentation is at docs.cactuscompute.com.
What industry is Cactus in?
Cactus's product category is On-device AI Inference Engine. Its primary akta.pro industry code is HDAAAJAB, On-Device Inference Runtimes & SDKs (mobile/embedded), with a secondary code of HDAAABAF, Model Deployment, Serving & Inference Platforms. Its NAICS code is 5132 and its SIC code is 7372.