Nexa AI
Nexa AI builds NexaSDK, a developer toolkit for running AI models locally across NPUs, GPUs, and CPUs on mobile, desktop, automotive, and IoT devices. Powered by its proprietary NexaML kernel-level inference engine, it targets privacy-preserving, offline-capable on-device AI. Nexa is now part of Qualcomm AI Hub.
- Company typePrivate
- Founded2023
- HeadquartersCupertino, United States
- Headcount11–50
- GTM typeB2B
- OfferingSoftware
What Nexa AI does
Nexa AI is a developer infrastructure company that builds NexaSDK, a unified toolkit for running AI models locally on device across mobile, desktop, automotive, and IoT hardware. The company's product surface covers five SDK targets — Android (Java/Kotlin via Maven Central), iOS/macOS (xcframework), Python (pip), a multi-OS CLI, and Docker for Linux ARM64 — and runs on NPUs, GPUs, and CPUs interchangeably, including Qualcomm Snapdragon (Hexagon NPU, Adreno GPU), Apple Silicon (ANE/MLX), and x86 CPUs. At the core is NexaML, a proprietary inference engine that Nexa describes as built from scratch at the kernel level rather than as a wrapper over an existing runtime, supporting three model formats (GGUF, MLX, and the proprietary .nexa format) and multiple AI modalities spanning text (LLM/VLM), image (CV/OCR/detection/depth/ImageGen), audio (ASR/TTS/diarization), embeddings, and reranking. Functional positioning emphasizes privacy-preserving, offline-capable, low-latency inference as alternatives to cloud-based AI APIs.
The business model is product-led and developer-first. Distribution is self-serve across public package registries, direct downloads, Docker Hub, GitHub, and Hugging Face, with developers onboarded through documentation and active Discord/Slack communities. The Builder Bounty Program pays up to $1,500 USD for developers who ship open-source applications on NexaSDK, and the Nexa Wishlist lets developers vote on which third-party models Nexa should optimize next. Monetization mechanics center on token-based SDK access for downloading NPU-optimized models, with GGUF and community models distributed without token gating; pricing is not publicly disclosed and there is no evidence of named enterprise contracts or disclosed ARR. Customer segmentation is horizontal and developer-led, targeting mobile app developers, desktop/PC integrators, embedded/IoT engineers, and automotive use cases.
Nexa AI is now operating as part of Qualcomm AI Hub following acquisition by Qualcomm Technologies, Inc., with the SDK redistributed through aihub.qualcomm.com. The company was founded in 2023, remains private, and runs with 11-50 employees out of Cupertino, California.
Nexa AI firmographics
Firmographics- Name
- Nexa AI
- Legal name
- Nexa AI
- Website
- https://nexa.ai
- Company type
- Private
- Founded year
- 2023
- Operating status
- Operating
- Headcount range
- 11–50 employees
- Short description
- Nexa AI builds NexaSDK, a developer toolkit for running AI models locally across NPUs, GPUs, and CPUs on mobile, desktop, automotive, and IoT devices. Powered by its proprietary NexaML kernel-level inference engine, it targets privacy-preserving, offline-capable on-device AI. Nexa is now part of Qualcomm AI Hub.
- Ownership category
- akta.pro rank
Nexa AI industry classification
Industry- Product category
- Edge AI Inference Software
- NAICS
- Software Publishers (5132), Software Publishers (513210)
- SIC
- Services-Prepackaged Software (7372)
- akta.pro primary industry
- On-Device Inference Runtimes & SDKs (mobile/embedded) (HDAAAJAB)
- akta.pro secondary industries
- AI Application Enablement Platforms (Copilot/Agent Frameworks, SDKs) (HDAEANAJ), Developer Tooling & SDK Platforms (BPAMAOAD), Platform Engineering & Internal Developer Platforms (IDP) (BPAEAKAC)
Keywords
Where Nexa AI is headquartered
LocationHeadquarters
- HQ city
- Cupertino
- HQ country
- United States
- HQ region
- North America
Markets served
Nexa AI business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Technology or R&D, Personnel, Infrastructure, Marketing or Sales
Revenue model
- SDK Access Token: Access tokens required to download and use NPU-optimized models. Users create accounts at sdk.nexa.ai and generate tokens via Deployment section.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Subscription | Pay-as-you-go | Developer SDK Access |
Go-to-market motion2 records
Distribution channels5 records
Marketing channels5 records
Nexa AI product offering
Product offeringCore offering
Nexa AI develops and distributes NexaSDK, a cross-platform developer toolkit that lets developers run AI models locally on NPUs, GPUs, and CPUs across Android, iOS/macOS, Windows, Linux, and automotive/IoT devices. The SDK is powered by the proprietary NexaML unified inference engine and exposes multimodal AI capabilities (LLM, VLM, ASR, TTS, Computer Vision, Embeddings, Reranking, ImageGen) via Python, CLI, REST API, and native mobile SDKs.
Product overview
Nexa AI offers NexaSDK, a unified developer toolkit for running AI models locally across hardware platforms (NPUs, GPUs, CPUs). The core is the NexaML inference engine supporting GGUF, MLX, and .nexa model formats. The product suite includes platform-specific SDKs for CLI/Go (Desktop), Python (Desktop), Android (mobile), iOS/macOS (mobile/desktop), and Docker (Linux/IoT). Key capabilities include LLM text generation, VLM multimodal processing, ASR/TTS speech processing, Embeddings/Reranker for search, Computer Vision (OCR, detection), and ImageGen. Nexa AI was acquired by Qualcomm and is now part of Qualcomm AI Hub.
Differentiator
Problem solved
Functional benefit
Brands
- NexaSDK: Developer toolkit for running AI models locally across NPUs, GPUs, and CPUs, powered by the NexaML engine
- NexaML
- Builder Bounty Program
Products and services
- NexaSDK A unified developer toolkit for running AI models locally across NPUs, GPUs, and CPUs, powered by the proprietary NexaML inference engine. Supports GGUF, MLX, and .nexa model formats with Day-0 support for new model architectures, and is available via Python, Android, iOS/macOS, Docker, and CLI for developers.
- Nexa CLI Command-line interface for downloading models (nexa pull), running inference (nexa infer), listing/removing models, and starting REST API server (nexa serve). Cross-platform support for macOS, Windows, and Linux.
- NexaSDK REST API Local OpenAI-compatible REST API for text generation, embeddings, image generation, and document reranking, with authentication via NEXA_TOKEN. Runs on port 18181 with Docker and CLI server modes for local deployment by developers.
- NexaAI Python SDK Python library for on-device AI inference supporting LLM, VLM, ASR, Embeddings, Reranker, CV, TTS, ImageGen, and Diarize modules. Supports macOS (MLX), Windows x64 (GGUF/CUDA), and Windows ARM64 (Snapdragon NPU) for Python developers.
- Nexa AI Android SDK Android SDK (v0.0.24) for on-device AI inference on Android devices, supporting LLM, VLM, Embeddings, ASR, Reranking, and Computer Vision with NPU (Qualcomm Hexagon), GPU (Adreno), and CPU acceleration. Minimum deployment target is ARM64-v8a for Android developers.
- Nexa AI iOS & macOS SDK iOS and macOS SDK (Beta) for on-device AI inference, supporting LLM, VLM, Embeddings, ASR, and Reranker models. Uses ANE (Apple Neural Engine) acceleration for ASR and Embeddings, GPU/CPU for other modules. Minimum: iOS 17.0, macOS 15.0.
- NexaSDK Docker Docker container solution for running NexaSDK on Linux ARM64 systems, optimized for Qualcomm NPU devices (Dragonwing IQ9) with support for server mode (REST API) and interactive CLI for embedded and IoT deployments.
Companies that use Nexa AI
Customer profileSegments4 records
Ideal customer profiles3 records
Nexa AI technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
AI capability10 records
Feature5 records
Nexa AI partnerships and signals
Strategic signalPartnerships
One partnership is on record.
- QualcommcoreNexa AI is now part of Qualcomm AI Hub. This represents a significant strategic acquisition/integration where Nexa AI's on-device AI SDK technology is incorporated into Qualcomm's AI ecosystem. Nexa SDK is now hosted and distributed through Qualcomm AI Hub at aihub.qualcomm.com.
Recent moves6 records
Expansion highlights6 records
Nexa AI competitors and assessment
Company assessmentDirect peers
- Ollama: Open-source tool for running large language models locally on consumer hardware. Directly comparable as an on-device/local LLM inference platform targeting developers, with overlapping model format support (GGUF) and cross-platform desktop coverage.
- llama.cpp: Open-source C/C++ inference engine for running LLMs locally on CPUs, GPUs, and various accelerators. Competes directly with NexaML as a low-level on-device inference engine, and is the de facto upstream project Nexa SDK supports via GGUF.
- MediaTek NeuroPilot: MediaTek's on-device AI SDK for Dimensity and other mobile SoCs. Directly comparable as an NPU-focused developer toolkit for on-device AI inference on mobile hardware, competing for the same embedded/edge developer audience.
Broad incumbents
- Apple Core ML: Apple's first-party on-device ML framework for iOS, macOS, and Apple Silicon. Overlaps with Nexa AI's iOS/macOS SDK in the same target audience (Apple developers) with a bundled, OS-integrated solution that competes on convenience and zero cost.
- Google ML Kit: Google's mobile SDK for on-device ML covering vision, text, and translation. Directly comparable to Nexa AI's Android SDK in target market (Android developers) and use cases (on-device CV, NLP, translation), with deep Google services integration.
- Qualcomm AI Engine SDK: Qualcomm's first-party Neural Processing SDK for Snapdragon. Now Nexa AI's parent company; previously a competitor offering NPU-targeted model conversion and inference — the acquisition effectively consolidated these capabilities under one roof.
- TensorFlow Lite (LiteRT): Google's lightweight inference runtime for mobile, embedded, and edge devices. Competes in the same on-device AI inference space across Android, iOS, and Linux/IoT — broader reach but higher complexity than Nexa SDK.
Emerging players
- Hugging Face: AI model hub and inference tooling with growing on-device offering (transformers.js, space support for local inference). Comparable as a developer platform where Nexa AI distributes GGUF and MLX models via Hugging Face collections.
- PyTorch Mobile: Meta's mobile deployment toolkit for PyTorch models. Overlaps with Nexa AI's mobile SDK as a developer-facing on-device ML deployment solution, though broader in scope and tied to the PyTorch ecosystem.
- ONNX Runtime: Cross-platform inference engine for ONNX models with mobile and edge support. Comparable as a low-level inference runtime targeting multiple hardware backends, competing with NexaML's kernel-level approach for model deployment.
Market position
Strengths4 records
Weaknesses4 records
Competitive moat4 records
Key risks5 records
Key highlights6 records
Customer concentration
Nexa AI social profiles
Digital presenceNexa AI financial estimates
Financial estimateRevenue estimate
Valuation estimate
Nexa AI leadership team
Management profileNumber of profiles
Nexa AI funding detail
Funding detailFunding overview
Funding rounds1 record
Investors3 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Nexa AI M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Nexa AI
What does Nexa AI do?
Nexa AI develops and distributes NexaSDK, a cross-platform developer toolkit that lets developers run AI models locally on NPUs, GPUs, and CPUs across Android, iOS/macOS, Windows, Linux, and automotive/IoT devices. The SDK is powered by the proprietary NexaML unified inference engine and exposes multimodal AI capabilities (LLM, VLM, ASR, TTS, Computer Vision, Embeddings, Reranking, ImageGen) via Python, CLI, REST API, and native mobile SDKs.
Is Nexa AI a public or private company?
Nexa AI is a private company. It is classified as corporate owned and is currently operating.
When was Nexa AI founded?
Nexa AI was founded in 2023. It employs 11 to 50 people.
Where is Nexa AI based?
Nexa AI is headquartered in Cupertino, United States, in the North America region.
How does Nexa AI make money?
One revenue line is on record: SDK Access Token.
Who are Nexa AI's main competitors?
Direct peers on record are Ollama, llama.cpp and MediaTek NeuroPilot. Broad incumbents are Apple Core ML, Google ML Kit, Qualcomm AI Engine SDK and TensorFlow Lite (LiteRT). Emerging players are Hugging Face, PyTorch Mobile and ONNX Runtime.
Does Nexa AI have an API?
Yes. NexaSDK provides a local OpenAI-compatible REST API for on-device AI inference. The API runs on http://127.0.0.1:18181 by default and supports multiple endpoints: /v1/chat/completions for LLM and VLM text generation, /v1/embeddings for text vectorization, /v1/images/generations for image synthesis, and /v1/reranking for document reranking. Authentication uses NEXA_TOKEN environment variable. The API is available via Docker containers for Linux ARM64 and x64, and supports CLI server mode for local deployment. Developer documentation is at docs.nexa.ai/en/nexa-sdk-go/NexaAPI.
What industry is Nexa AI in?
Nexa AI's product category is Edge AI Inference Software. Its primary akta.pro industry code is HDAAAJAB, On-Device Inference Runtimes & SDKs (mobile/embedded), with a secondary code of HDAEANAJ, AI Application Enablement Platforms (Copilot/Agent Frameworks, SDKs). Its NAICS code is 5132 and its SIC code is 7372.