Neural Magic
Neural Magic provides AI inference optimization software for deploying open-source large language models, using SparseGPT compression and the llm-d distributed inference framework. Acquired by Red Hat in January 2025; serves enterprise hybrid cloud AI customers.
- Company typePrivate
- Founded2018
- HeadquartersSomerville, United States
- Headcount51–100
- GTM typeB2B
- OfferingSoftware
What Neural Magic does
Neural Magic, founded in 2018 and headquartered in Somerville, Massachusetts, built software for accelerating generative AI inference workloads, with a focus on open-source large language model deployment. Its core technology stack applies model compression techniques — most notably SparseGPT one-shot pruning, quantization, and knowledge distillation — to reduce model size and computational requirements while preserving accuracy, allowing LLMs to run efficiently across heterogeneous hardware including NVIDIA and AMD GPUs and Google TPUs. The company authored the llm-d open-source distributed inference framework, designed as a production-grade Kubernetes workload for vendor-neutral, scalable LLM serving, and integrated with the vLLM serving engine.
The company's commercial trajectory culminated in its January 2025 acquisition by Red Hat (an IBM subsidiary). Prior to the exit, Neural Magic had raised approximately $50 million across seed, Series 1 (Comcast Ventures-led, 2019), and Series A (NEA-led, 2021) rounds from investors including Andreessen Horowitz, New Enterprise Associates, Comcast Ventures, Pillar VC, Amdocs, and Ridgeline. Post-acquisition, Neural Magic's compression tooling is embedded inside Red Hat AI Inference Server, RHEL AI, and OpenShift AI, and its llm-d framework was donated to the Cloud Native Computing Foundation as a sandbox project in March 2026 with ten founding industry supporters (NVIDIA, CoreWeave, AMD, Cisco, Hugging Face, Intel, Lambda, Mistral AI, IBM Research, and Google Cloud). The company's go-to-market is now channel-driven through Red Hat's enterprise salesforce and open-source community distribution, with no independently disclosed pricing or standalone revenue figures.
Neural Magic firmographics
Firmographics- Name
- Neural Magic
- Legal name
- Neural Magic Inc.
- Website
- https://neuralmagic.com
- Company type
- Private
- Founded year
- 2018
- Operating status
- Acquired
- Headcount range
- 51–100 employees
- Short description
- Neural Magic provides AI inference optimization software for deploying open-source large language models, using SparseGPT compression and the llm-d distributed inference framework. Acquired by Red Hat in January 2025; serves enterprise hybrid cloud AI customers.
- Ownership category
- akta.pro rank
Neural Magic industry classification
Industry- Product category
- AI Inference Optimization Software
- NAICS
- Software Publishers (5132), Computer Systems Design and Related Services (5415), Custom Computer Programming Services (541511)
- SIC
- Services-Prepackaged Software (7372), Services-Computer Integrated Systems Design (7373)
- akta.pro primary industry
- Model Compression & Optimization (Quantization, Distillation, Pruning) (HDAAACAL)
- akta.pro secondary industries
- Edge AI Model Optimization & Compression (quantization, pruning, distillation) (HDAAAJAA), LLMOps & Generative AI Platforms (Prompt/Agent Orchestration, RAG) (HDAEANAD)
Keywords
Where Neural Magic is headquartered
LocationHeadquarters
- HQ city
- Somerville
- HQ country
- United States
- HQ region
- North America
Markets served
Neural Magic business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Technology or R&D, Personnel, Infrastructure, Marketing or Sales, Operations
Revenue model
- Technology Licensing to Red Hat: Neural Magic's technology was acquired by Red Hat in 2025, providing model optimization and performance engineering capabilities for Red Hat's AI platform focused on hybrid cloud environments.
Go-to-market motion2 records
Distribution channels3 records
Marketing channels3 records
Neural Magic product offering
Product offeringCore offering
Neural Magic develops AI inference optimization software that uses model compression techniques — pruning, quantization, knowledge distillation, and the proprietary SparseGPT algorithm — to run large language models efficiently on commodity and accelerated hardware. Its llm-d framework provides distributed, Kubernetes-native LLM serving, now integrated into Red Hat's AI platform (Red Hat AI Inference Server, OpenShift AI, RHEL AI) following Red Hat's January 2025 acquisition.
Product overview
Neural Magic was an independent company specializing in software and algorithms for accelerating genAI inference workloads, focusing on model optimization, performance engineering, and open-source AI solutions. The company developed the llm-d distributed inference framework for Kubernetes and model compression tools using SparseGPT techniques. Neural Magic was acquired by Red Hat in January 2025, and its technology now integrates into Red Hat's AI platform portfolio including Red Hat AI Inference Server and RHEL AI, providing enterprises with standardized open-source building blocks for AI workloads.
Differentiator
Problem solved
Functional benefit
Products and services
- llm-d open-source distributed inference framework Open-source distributed inference framework for Kubernetes that enables scalable, vendor-neutral LLM inference as a production-grade cloud-native workload. Created by Neural Magic and donated to CNCF as a sandbox project, supported by NVIDIA, AMD, Intel, Hugging Face, and other founding partners.
- SparseGPT model compression tools Proprietary model compression tools developed by Neural Magic using the SparseGPT technique, now integrated into Red Hat AI Inference Server. Reduce LLM model size and computational requirements through pruning and quantization while preserving accuracy.
- Sparse LLM training and deployment capabilities (with Cerebras) Joint technology offering with Cerebras for training and deploying sparse large language models, designed to deliver faster, more power-efficient, and lower-cost AI model training and deployment versus dense-model approaches.
Quantifiable outcome
- 2x improvements in time-to-first-token for code completion use cases (Google Cloud testing)
Companies that use Neural Magic
Customer profileIdeal customer profiles2 records
Neural Magic technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- No
- API docs
- API detail
Core technology
AI maturity
App detail
Integration1 record
AI capability4 records
Feature4 records
Neural Magic partnerships and signals
Strategic signalPartnerships
Nine partnerships are on record, tiered founding and core.
- NVIDIAfoundingNVIDIA provided founding support for the llm-d open-source distributed inference framework donated to CNCF. NVIDIA joins other founding supporters including CoreWeave, AMD, Cisco, Hugging Face, Intel, Lambda, and Mistral AI.
- CoreWeavefoundingCoreWeave provided founding support for the llm-d open-source distributed inference framework donated to CNCF.
- AMDfoundingAMD provided founding support for the llm-d open-source distributed inference framework donated to CNCF.
- CiscofoundingCisco provided founding support for the llm-d open-source distributed inference framework donated to CNCF.
- Hugging FacefoundingHugging Face provided founding support for the llm-d open-source distributed inference framework donated to CNCF.
- IntelfoundingIntel provided founding support for the llm-d open-source distributed inference framework donated to CNCF.
- LambdafoundingLambda provided founding support for the llm-d open-source distributed inference framework donated to CNCF.
- Mistral AIfoundingMistral AI provided founding support for the llm-d open-source distributed inference framework donated to CNCF.
- CerebrascoreCerebras and Neural Magic jointly announced new capabilities for training and deploying sparse large language models. The partnership aims to address efficiency challenges in AI model training and deployment, positioning the technology as faster, more power efficient, and lower cost.
Scale indicators3 records
Recent moves6 records
Expansion highlights5 records
Neural Magic competitors and assessment
Company assessmentDirect peers
- Hugging Face: Hugging Face is a founding supporter of the llm-d framework and operates the leading hub for open-source LLMs along with its own inference and optimization tooling (TGI, Text Embeddings Inference). It directly competes in open-source model deployment and inference optimization for enterprise customers.
- vLLM Project: vLLM is the open-source high-throughput LLM serving engine integrated into Red Hat AI Inference Server alongside Neural Magic's compression tools. It is a direct peer as the primary inference runtime that Neural Magic's optimizations are paired with, with overlapping contributors and roadmap.
- Anyscale: Anyscale (creators of Ray and Ray Serve) provides scalable open-source AI compute and inference infrastructure for production LLM workloads. It competes with Neural Magic/llm-d as a Kubernetes-native, scalable inference platform for enterprise generative AI.
- Together AI: Together AI offers an open-source-focused AI inference cloud with optimized inference engines and model compression capabilities for serving LLMs at scale. It competes directly with Neural Magic in optimized, cost-efficient open-source LLM inference for enterprise customers.
- Fireworks AI: Fireworks AI provides a managed inference platform with proprietary optimizations for open-source LLMs, focusing on low-latency, high-throughput production deployment. It is a direct competitor in enterprise-grade LLM inference optimization.
- DeepInfra: DeepInfra provides low-cost, low-latency hosted inference for open-source LLMs using proprietary optimizations. It competes directly with Neural Magic's positioning around cost-efficient, high-performance open-source LLM inference for enterprise and developer customers.
Broad incumbents
- NVIDIA (TensorRT-LLM): NVIDIA's TensorRT-LLM is the dominant proprietary inference optimization stack tightly coupled to NVIDIA GPUs, with extensive kernel and quantization support. While Neural Magic supports NVIDIA hardware, NVIDIA competes broadly in the same inference optimization layer with deeper hardware-specific engineering.
Emerging players
- Cerebras Systems: Cerebras is a strategic partner of Neural Magic for sparse LLM training and deployment, and also competes in the inference optimization space via its wafer-scale accelerator stack and dedicated inference solutions. Highly relevant due to the joint go-to-market and overlapping inference efficiency value proposition.
- Modular: Modular provides the Mojo language and MAX inference platform for high-performance AI deployment across heterogeneous hardware. It competes in the same inference optimization and hardware-agnostic deployment category as Neural Magic's llm-d stack.
- OctoAI (now part of NVIDIA): OctoAI built a managed inference optimization platform for open-source models before being acquired by NVIDIA. It is a comparable inference optimization player for production LLM workloads and now sits inside NVIDIA's broader stack.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat4 records
Key risks5 records
Key highlights6 records
Customer concentration
Neural Magic social profiles
Digital presenceNeural Magic financial estimates
Financial estimateRevenue estimate
Valuation estimate
Neural Magic leadership team
Management profileNumber of profiles
Profiles4 records
Neural Magic funding detail
Funding detailFunding overview
Funding rounds3 records
Investors7 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Neural Magic M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Neural Magic
What does Neural Magic do?
Neural Magic develops AI inference optimization software that uses model compression techniques — pruning, quantization, knowledge distillation, and the proprietary SparseGPT algorithm — to run large language models efficiently on commodity and accelerated hardware. Its llm-d framework provides distributed, Kubernetes-native LLM serving, now integrated into Red Hat's AI platform (Red Hat AI Inference Server, OpenShift AI, RHEL AI) following Red Hat's January 2025 acquisition.
Is Neural Magic a public or private company?
Neural Magic is a private company. It is classified as corporate owned and is currently acquired.
When was Neural Magic founded?
Neural Magic was founded in 2018. It employs 51 to 100 people.
Where is Neural Magic based?
Neural Magic is headquartered in Somerville, United States, in the North America region.
How does Neural Magic make money?
One revenue line is on record: technology Licensing to Red Hat.
Who are Neural Magic's main competitors?
Direct peers on record are Hugging Face, vLLM Project, Anyscale, Together AI, Fireworks AI and DeepInfra. NVIDIA (TensorRT-LLM) is listed as a broad incumbent. Emerging players are Cerebras Systems, Modular and OctoAI (now part of NVIDIA).
Does Neural Magic have an API?
No public API is recorded for Neural Magic.
What industry is Neural Magic in?
Neural Magic's product category is AI Inference Optimization Software. Its primary akta.pro industry code is HDAAACAL, Model Compression & Optimization (Quantization, Distillation, Pruning), with a secondary code of HDAAAJAA, Edge AI Model Optimization & Compression (quantization, pruning, distillation). Its NAICS code is 5132 and its SIC code is 7372.