Moreh
Moreh is a Seoul-based AI infrastructure software company building the MoAI full-stack inference platform that orchestrates heterogeneous GPU clusters (AMD, NVIDIA, Tenstorrent), serving AI data centers, cloud providers, telecom operators, and government AI initiatives with cross-vendor disaggregation and custom inference engines.
- Company typePrivate
- Founded2020
- HeadquartersSeoul, South Korea
- Headcount51–100
- GTM typeB2B
- OfferingSoftware
What Moreh does
Moreh is an AI infrastructure software company founded in September 2020 in Seoul, South Korea, that builds a full-stack inference platform — the MoAI framework — optimized for heterogeneous GPU clusters spanning AMD Instinct, NVIDIA, and Tenstorrent accelerators. The core product line includes the MoAI Inference Framework (cluster-scale routing, scheduling, auto-scaling, KV cache management), Moreh vLLM drop-in replacements for AMD and Tenstorrent GPUs, MoAI Performance Gateway for cross-vendor workload distribution, MoAI Fabric for RDMA-based KV cache transfer across vendors, and proprietary primitives including HetCCL (cross-vendor collective communication), SLOPE (long-context prefill engine), and TIDE (runtime draft model trainer). The company reports differentiated capabilities such as cross-vendor prefill-decode disaggregation between NVIDIA H100 and AMD MI300X, with published benchmarks claiming up to 1.84x lower latency and 67% higher throughput versus single-vendor baselines.
The company monetizes via enterprise software licensing (subscription, quote-based, multi-year) combined with professional and managed services that bundle hardware procurement, cluster design and deployment, and ongoing optimization for AMD and Tenstorrent clusters; pricing is not publicly disclosed. Distribution runs through direct enterprise sales to AI data centers, cloud service providers, and telecom operators, augmented by strategic channel partnerships with AMD and Tenstorrent for joint go-to-market and open-source contributions to ROCm, SGLang, SkyPilot, and Tenstorrent Metalium. Key customers include KT Cloud (Hyperscale AI Computing service running 1,200+ AMD MI250 GPUs) and the Korea Ministry of Science and ICT (lead consortium for the national sovereign AI foundation model project); subsidiary Motif Technologies develops proprietary Korean-language LLMs. The company is backed by AMD, KT, and Forest Ventures, with cumulative funding exceeding $50 million USD and a post-money valuation of approximately 350 billion KRW (~$250M USD) as of September 2025.
Moreh firmographics
Firmographics- Name
- Moreh
- Legal name
- Moreh, Inc.
- Website
- https://moreh.io
- Company type
- Private
- Founded year
- 2020
- Operating status
- Operating
- Headcount range
- 51–100 employees
- Short description
- Moreh is a Seoul-based AI infrastructure software company building the MoAI full-stack inference platform that orchestrates heterogeneous GPU clusters (AMD, NVIDIA, Tenstorrent), serving AI data centers, cloud providers, telecom operators, and government AI initiatives with cross-vendor disaggregation and custom inference engines.
- Ownership category
- akta.pro rank
Moreh industry classification
Industry- Product category
- AI Inference Software
- NAICS
- Computer Systems Design and Related Services (5415), Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (518), Custom Computer Programming Services (541511)
- SIC
- Services-Prepackaged Software (7372), Services-Computer Integrated Systems Design (7373), Services-Computer Programming, Data Processing, Etc. (7370)
- akta.pro primary industry
- Model Serving, Inference & Deployment Platforms (APIs, Edge/On-Prem) (HDAEANAC)
- akta.pro secondary industries
- AI Compiler, Runtime & Kernel Optimization Software (CUDA/ROCm/XLA, graph compilers) (HDAAAAAI), AI Compute Virtualization & Scheduling (GPU virtualization, cluster schedulers) (HDAAAAAG), Model Deployment, Serving & Inference Platforms (HDAAABAF), Model Hosting, Serving & Inference Platforms (HDAAACAB)
Keywords
Where Moreh is headquartered
LocationHeadquarters
- HQ city
- Seoul
- HQ country
- South Korea
- HQ region
- Asia
Offices3 records
Markets served
Moreh business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Personnel, Technology or R&D, Operations, Marketing or Sales, Infrastructure
Revenue model
- AI Infrastructure Software Licensing: Software licensing for AI inference infrastructure, including MoAI Inference Framework and Moreh vLLM. Provided as container images with regular updates for AMD GPU clusters and Tenstorrent solutions. Pricing appears to be enterprise quote-based with multi-year contracts.
- Turnkey GPU Cluster Solutions: End-to-end AMD GPU cluster solutions combining hardware procurement, cluster design/building, software deployment, and ongoing technical support. Covers rack layout, power planning, network topology design, and deployment of Moreh inference software.
- Tenstorrent AI Cluster Solutions: Scalable AI cluster solutions built around Tenstorrent Wormhole processors, including hardware supply, cluster construction, software deployment, and continuous technical support for Tenstorrent-specific issues and performance tuning.
- Custom Benchmarking Services: Custom benchmarking engagements where Moreh runs performance tests on customer workloads to demonstrate cost reduction effects and token-per-dollar improvements, likely as part of enterprise sales process.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Subscription | Multi-year contract | Enterprise licensing with custom benchmarking |
Go-to-market motion3 records
Distribution channels3 records
Marketing channels6 records
Moreh product offering
Product offeringCore offering
Moreh develops and licenses full-stack AI inference software (the MoAI platform) that enables organizations to run and serve large language models on heterogeneous GPU clusters combining AMD Instinct, NVIDIA, and Tenstorrent accelerators. The offering spans chip-level custom kernels, Moreh vLLM inference engines optimized for AMD and Tenstorrent chips, cluster-level orchestration with cross-vendor prefill-decode disaggregation, and turnkey GPU cluster solutions for AI data centers, cloud service providers, telecom operators, and government AI projects.
Product overview
Moreh is an AI infrastructure software company founded in 2020 by HPC experts, offering a modular platform-plus-modules architecture. Its core products include the MoAI Inference Framework (cluster-scale orchestration), Moreh vLLM for AMD and Tenstorrent (chip-level inference engines), MoAI Performance Gateway (intelligent workload routing), MoAI Fabric (cross-vendor KV cache transfer), and MoAI Training Framework (fine-tuning and training). Additional capabilities include SLOPE (long-context prefill engine), TIDE (runtime draft model training), and HetCCL (cross-vendor collective communication). Moreh delivers turnkey GPU cluster solutions for AMD Instinct and Tenstorrent hardware, enabling organizations to unify heterogeneous GPUs across vendors and generations into a single inference cluster. Its subsidiary Motif Technologies develops Korean-language generative AI models. The company targets AI data centers seeking to maximize tokens per dollar through chip-level, cluster-level, and infrastructure optimization.
Differentiator
Problem solved
Functional benefit
Products and services
- MoAI Inference Framework Production-grade end-to-end inference stack optimized for heterogeneous accelerators (AMD, NVIDIA, Tenstorrent) that provides routing and scheduling, auto scaling, SLO-driven optimization, and KV cache management across cluster-scale deployments. Targets enterprise AI data centers running large language models.
- MoAI Performance Gateway Intelligent workload distribution component across heterogeneous accelerators, enabling cross-vendor GPU orchestration and automated optimization decisions within the MoAI Inference Framework. Provides dynamic routing and load balancing across GPUs from different vendors and generations.
- MoAI Fabric Software-defined, cross-vendor GPU memory fabric for KV cache transfer, enabling RDMA-based communication between GPUs from different vendors (e.g., NVIDIA and AMD) across cluster nodes for inference workloads.
- Moreh vLLM for AMD Drop-in vLLM replacement optimized for AMD GPUs that delivers up to 2x higher throughput on AMD Instinct GPUs. Includes custom kernels for GEMM, Attention, and MoE operations, and is distributed as container images with regular updates for AMD GPU clusters.
- Moreh vLLM for Tenstorrent High-performance vLLM serving engine purpose-built for Tenstorrent Wormhole accelerators. Supports state-of-the-art MoE models including DeepSeek, GPT-OSS, and Qwen, with 450+ optimized operators for general heterogeneous GPU inference.
- MoAI Training Framework Training and fine-tuning framework for AMD and Tenstorrent GPU clusters. PyTorch-compatible with 450+ operators, enabling fine-tuning and training on the same cluster used for inference.
- Turnkey AMD GPU Clusters Full-stack turnkey AMD Instinct GPU cluster solution that includes hardware procurement, rack layout and power planning, network topology design (RoCE), Kubernetes platform deployment, and Moreh's inference software stack. Delivered to AI data centers, with 1,200+ AMD MI250 GPUs deployed at KT Cloud for training operations.
- Tenstorrent AI Clusters Cost-efficient AI cluster solution built around Tenstorrent's network-integrated Wormhole processors and Galaxy servers. Moreh provides end-to-end support from hardware supply and cluster building to software deployment and ongoing technical support for Tenstorrent-specific issues and performance tuning.
- Heterogeneous GPU Inference Solution Solution unifying GPUs across vendors, architectures, and generations into a single inference cluster. Supports three scenarios: legacy + new generation GPUs (e.g., H100 + B200), NVIDIA + AMD mixed clusters (e.g., H200 + MI355X), and GPU + AI accelerator mixed (e.g., GPU + Tenstorrent).
- Inference Cost Optimization Solution Cost-per-token optimization solution operating across three levers: chip-level optimization (Moreh vLLM), cluster-level optimization (PD disaggregation, prefix caching, auto scaling), and infrastructure cost reduction (heterogeneous GPU utilization). Achieves up to 2.2x throughput on 40% fewer servers.
- Motif Technologies AI Models Subsidiary that develops Korean-language generative AI models including Motif 12.7B with proprietary GDA (group differential attention) architecture, plus cloud services. Selected as lead team for Korea's government 'Independent AI Foundation Model' project developing multimodal foundation models across text, images, video, and audio.
Quantifiable outcome
- 1.68x throughput improvement for DeepSeek R1 671B on AMD MI300X vs ROCm vLLM baseline
- +6 more outcomes
Companies that use Moreh
Customer profileNamed customers4 records
Segments4 records
Ideal customer profiles4 records
Moreh technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration4 records
AI capability6 records
Feature10 records
Moreh partnerships and signals
Strategic signalPartnerships
Five partnerships are on record, tiered core and secondary.
- Korea Ministry of Science and ICTcoreSelected as lead team for Korea's 'Independent AI Foundation Model' project, forming consortium with 17 participating institutions and 12 demand-side institutions. Consortium includes Moreh for GPU optimization, academic institutions for multimodal design, and companies for data creation and synthetic VLA data generation.
- SGLangsecondaryCollaboration to expand into deep learning inference market. Joint development of AMD-based distributed inference system. Showcase of distributed inference system on AMD hardware at AI Infra Summit 2025, demonstrating DeepSeek model optimization more efficient than NVIDIA systems.
- Motif TechnologiescoreSubsidiary established in February 2025 for model development and cloud services. Led by Motif CEO Lim Jung-hwan, developing proprietary AI architecture including GDA (group differential attention) structure. Motif 12.7B model ranked first among domestic models in Artificial Analysis intelligence index. Selected as lead team for Korea's sovereign AI foundation model project.
- TenstorrentcoreStrategic partnership to challenge NVIDIA in AI data center market. Combines Tenstorrent's AI semiconductors (NPUs/Wormhole processors) with Moreh's parallel processing software to support inference and LLM training applications. Planned for commercialization in H1 2025. Joint AI data center solution unveiled at SuperComputing 2025 with cost-efficient alternative to NVIDIA infrastructure.
- KT CloudcoreLaunched Hyperscale AI Computing service in December 2021 based on Moreh's MoAI platform and AMD Instinct MI250 accelerators. Service became paid in August 2022. Selected as official cloud service for Korean government HPC program. KT Cloud resells GPU access through Moreh's virtualization platform.
Scale indicators13 records
Recent moves10 records
Expansion highlights6 records
Moreh competitors and assessment
Company assessmentDirect peers
- Anyscale: Anyscale operates a Ray-based AI compute platform for production AI workloads including LLM serving and batch inference. Comparable to Moreh as a platform layer that abstracts heterogeneous compute (Ray clusters vs Moreh's MoAI cluster abstraction) for enterprise AI deployment, with similar developer-led go-to-market and OSS roots.
- Together AI: Together AI provides an AI cloud and inference platform optimized for open-source LLMs with custom kernels and inference engines. Directly comparable to Moreh in offering inference-as-a-service with proprietary optimizations on top of open-source serving frameworks (vLLM, FlashAttention), targeting cost-sensitive enterprise and developer customers.
- Fireworks AI: Fireworks AI is a model-serving platform optimized for low-latency LLM inference with proprietary kernel and routing optimizations. Comparable to Moreh's MoAI Inference Framework as a production-grade inference orchestration layer, with similar focus on throughput-per-dollar and serving reliability for frontier models.
- Modular: Modular builds the Mojo language and a unified AI inference platform (MAX) that targets heterogeneous accelerators with high-performance compilation. Comparable to Moreh in pursuing compiler/runtime-level optimization (Moreh IR vs Modular's Mojo/MAX) and in abstracting GPUs from the application developer, with similar AI infrastructure software positioning.
- OctoAI (acquired by NVIDIA): OctoAI built an enterprise AI inference platform optimized across hardware backends before being acquired by NVIDIA. Comparable as a prior independent inference-software specialist competing on cost-per-token across heterogeneous accelerators — a direct competitive template Moreh continues to operate within.
- Replicate: Replicate operates a cloud platform for running and deploying ML models, with proprietary inference optimization on heterogeneous hardware (AMD, NVIDIA, Apple Silicon). Comparable to Moreh's MoAI as a software layer that makes running open-source models efficient across accelerators, targeting developers and enterprises.
Emerging players
- RunPod: RunPod provides GPU cloud infrastructure with inference endpoints and custom kernel optimizations for AMD and NVIDIA hardware. Comparable as a smaller, AMD-friendly alternative GPU cloud whose software stack overlaps with Moreh's turnkey AMD cluster offerings for cost-sensitive AI customers.
Broad incumbents
- CoreWeave: CoreWeave is a large-scale GPU cloud provider with proprietary orchestration and inference serving software (TensorWave). Comparable as an AI infrastructure operator competing for the same cost-optimized inference workloads, though NVIDIA-centric and significantly larger and more capitalized than Moreh.
- Lambda Labs: Lambda provides GPU cloud, on-prem clusters, and inference serving software for AI workloads. Comparable to Moreh in offering turnkey GPU clusters with proprietary software optimizations, though primarily NVIDIA-focused and operating at a larger scale.
Others
- Tenstorrent: Tenstorrent is an AI accelerator (NPU) company and Moreh's strategic partner, with joint AI data center solutions unveiled at SuperComputing 2025. Listed as a peer rather than partner because Tenstorrent's success directly determines Moreh's addressable market on the Tenstorrent hardware ecosystem; the strategic alignment is also a mutual competitive dependency.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat5 records
Key risks6 records
Key highlights7 records
Customer concentration
Moreh social profiles
Digital presenceMoreh compliance and trust
Trust signalCompliance1 record
Moreh financial estimates
Financial estimateRevenue estimate
Valuation estimate
Moreh leadership team
Management profileNumber of profiles
Profiles6 records
Moreh subsidiaries and ownership
Company hierarchySubsidiaries2 records
Moreh funding detail
Funding detailFunding overview
Funding rounds6 records
Investors4 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Moreh M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Moreh
What does Moreh do?
Moreh develops and licenses full-stack AI inference software (the MoAI platform) that enables organizations to run and serve large language models on heterogeneous GPU clusters combining AMD Instinct, NVIDIA, and Tenstorrent accelerators. The offering spans chip-level custom kernels, Moreh vLLM inference engines optimized for AMD and Tenstorrent chips, cluster-level orchestration with cross-vendor prefill-decode disaggregation, and turnkey GPU cluster solutions for AI data centers, cloud service providers, telecom operators, and government AI projects.
Is Moreh a public or private company?
Moreh is a private company. It is classified as venture growth investor backed and is currently operating.
When was Moreh founded?
Moreh was founded in 2020. It employs 51 to 100 people.
Where is Moreh based?
Moreh is headquartered in Seoul, South Korea, in the Asia region.
How does Moreh make money?
Four revenue lines are on record. AI Infrastructure Software Licensing is the primary driver. The others are turnkey GPU Cluster Solutions, tenstorrent AI Cluster Solutions and custom Benchmarking Services.
Who are Moreh's main competitors?
Direct peers on record are Anyscale, Together AI, Fireworks AI, Modular, OctoAI (acquired by NVIDIA) and Replicate. RunPod is listed as an emerging player. Broad incumbents are CoreWeave and Lambda Labs. Tenstorrent is listed as an others.
Does Moreh have an API?
Yes. Moreh provides an OpenAI-compatible API endpoint across its cluster deployments via the MoAI Inference Framework, allowing developers to route inference requests through a unified cluster-wide API. The API supports both inference and training workloads and is accessible via standard REST calls. Developer documentation is at docs.moreh.io.
What industry is Moreh in?
Moreh's product category is AI Inference Software. Its primary akta.pro industry code is HDAEANAC, Model Serving, Inference & Deployment Platforms (APIs, Edge/On-Prem), with a secondary code of HDAAAAAI, AI Compiler, Runtime & Kernel Optimization Software (CUDA/ROCm/XLA, graph compilers). Its NAICS code is 5415 and its SIC code is 7372.