Zyphra
Zyphra is an open superintelligence company building open-weight foundation models and a full-stack AMD-native AI cloud platform serving developers, enterprises, and frontier AI hyperscalers.
- Company typePrivate
- Founded2021
- HeadquartersPalo Alto, United States
- Headcount51–100
- GTM typeB2B
- OfferingSoftware
What Zyphra does
Zyphra is an open superintelligence research and product company headquartered at 415 Mission St, Floor 44 (Salesforce Tower), San Francisco, with an additional operations and hiring presence in London. The company operates two interlocking divisions: Zyphra Research, which develops novel foundation model architectures and open-weight models, and Zyphra Cloud, a full-stack AI platform built exclusively on AMD Instinct GPUs (MI300X for training, MI355X for production inference via TensorWave infrastructure). Its model portfolio spans language (ZAYA1-8B, ZAYA1-74B-Preview, Zamba2 series, ZR1-1.5B), vision-language (ZAYA1-VL-8B, Zamba2-VL), text-to-speech (Zonos/ZONOS2), BCI/EEG (ZUNA), and diffusion-language (ZAYA1-8B-Diffusion-Preview), all released under Apache 2.0. Core technical differentiators include Compressed Convolutional Attention (8x KV-cache compression), Tensor and Sequence Parallelism (2.6x throughput at 1,024 GPU scale), Tree Attention decoding, MLP-based expert routers with PID-style bias balancing, Markovian RSA test-time compute, and custom AMD ROCm kernels.
Zyphra firmographics
Firmographics- Name
- Zyphra
- Legal name
- Zyphra Technologies, Inc.
- Website
- https://zyphra.com
- Company type
- Private
- Founded year
- 2021
- Operating status
- Operating
- Headcount range
- 51–100 employees
- Short description
- Zyphra is an open superintelligence company building open-weight foundation models and a full-stack AMD-native AI cloud platform serving developers, enterprises, and frontier AI hyperscalers.
- Ownership category
- akta.pro rank
Zyphra industry classification
Industry- Product category
- AI Foundation Models & Cloud Platform
- NAICS
- Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (5182)
- SIC
- Services-Prepackaged Software (7372)
- akta.pro primary industry
- Model Hosting, Serving & Inference Platforms (HDAAACAB)
- akta.pro secondary industries
- AI Compiler, Runtime & Kernel Optimization Software (CUDA/ROCm/XLA, graph compilers) (HDAAAAAI), Open-Source Model Ecosystems & Model Marketplaces (HDAAACAM), Model Compression & Optimization (Quantization, Distillation, Pruning) (HDAAACAL), Decentralized GPU/AI Compute Networks (training/inference, GPU renting, model hosting) (FSAPALAC)
Keywords
Where Zyphra is headquartered
LocationHeadquarters
- HQ city
- Palo Alto
- HQ country
- United States
- HQ region
- North America
Offices3 records
Markets served
Zyphra business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Technology or R&D, Personnel, Infrastructure, Operations, Marketing or Sales
Revenue model
- Zyphra Cloud Serverless Inference: Usage-based pay-as-you-go inference for frontier open-weight models including Kimi K2.6, DeepSeek V3.2, and GLM 5.1 on AMD MI355X GPUs; available via cloud.zyphra.com with signup
- Enterprise Dedicated Capacity and Custom Hyperscale Infrastructure: Reserved capacity for latency-sensitive production deployments and custom hyperscale AMD bare-metal GPU cluster deployments for large-scale training and inference; sales-led contract revenue
- Agent Platform and Compute Services: Future-revenue streams include MAIA agent platform, agent environments, distributed training, RL, and large-scale simulation environments running on Zyphra Cloud
- Promotional Free Service (ZONOS2 launch): ZONOS2 served free of charge for a promotional period after release to drive adoption and developer mindshare
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Usage-based | Pay-as-you-go | Serverless inference for frontier open-weight models on AMD MI355X |
| Subscription | Multi-year contract | Reserved / dedicated capacity for latency-sensitive production |
| Freemium | Pay-as-you-go | Free serverless endpoint for ZAYA1-8B |
| Freemium | Pay-as-you-go | Promotional free access to ZONOS2 |
Go-to-market motion1 record
Distribution channels5 records
Marketing channels8 records
Zyphra product offering
Product offeringCore offering
Zyphra operates a full-stack open AI platform combining Zyphra Research (novel foundation model architectures including MoE with CCA attention and custom AMD-native training/inference kernels) and Zyphra Cloud (bare-metal AMD GPU infrastructure delivering serverless and dedicated inference for frontier open-weight models, the MAIA multiplayer agent platform, and custom hyperscale deployments). The company releases open-weight Apache 2.0 foundation models across language, vision-language, text-to-speech, diffusion, and BCI modalities, with proprietary parallelism and decoding schemes optimized for AMD MI300X/MI355X hardware.
Product overview
Zyphra operates a platform-plus-modules architecture split between two divisions: Zyphra Research develops open foundation models and novel architectures, while Zyphra Cloud productizes those innovations as a full-stack AI platform. The core platform is Zyphra Cloud, which launched with Zyphra Inference (serverless inference for frontier open-weight models like DeepSeek V3.2, Kimi K2.6, and GLM 5.1, plus Zyphra's own models), bare-metal AMD GPU capacity, the MAIA multiplayer superagent, and agent environments. The Zyphra Research side contributes the ZAYA family of language/vision-language/diffusion models (ZAYA1-8B, ZAYA1-74B-Preview, ZAYA1-VL-8B, ZAYA1-8B-Diffusion-Preview), the Zonos/ZONOS2 TTS models, the ZUNA BCI foundation model, the Zamba2 series of small hybrid SSM-Transformer language models, ZR1-1.5B reasoning model, and datasets (Zyda, Zyda-2). Together, Zyphra Research and Zyphra Cloud form a continuous cycle of discovery and delivery for open superintelligence.
Differentiator
Problem solved
Functional benefit
Brands
- Zyphra Research: Zyphra's research lab focused on next-generation model architectures, multimodal world models, and silicon performance; produces open foundation models such as ZAYA, ZONOS, and ZUNA.
- Zyphra Cloud
- Zyphra Inference
- MAIA
- ZAYA
- ZONOS
- ZUNA
- ZAYA-VL
Products and services
- Zyphra Cloud Full-stack AI platform on AMD Instinct MI355X GPUs (TensorWave infrastructure) delivering Zyphra Inference (serverless and dedicated), the MAIA agent, agent environments (CPUs), distributed training & RL environments, and bare-metal GPU cluster orchestration for developers, enterprises, and frontier AI hyperscalers.
- Zyphra Inference AMD-first inference service purpose-built for large MoE models and long-running agentic workloads with long context and large KV/prefix caches. Serves frontier open-weight models (Kimi K2.6, DeepSeek V3.2, GLM 5.1) plus Zyphra's own models via serverless endpoints and dedicated inference capacity.
- MAIA Multiplayer general-purpose open superagent for teams. Unified multimodal system that coordinates knowledge, communication, and execution across tools and workflows with shared context, persistent memory, and language/audio/vision interaction in a single reasoning loop.
- ZAYA1-8B Open MoE reasoning language model with under 1B active parameters (~8.4B total), pretrained end-to-end on AMD MI300X clusters, incorporating Compressed Convolutional Attention (CCA), a novel MLP-based router, learned residual scaling, and the Markovian RSA test-time compute methodology. Apache 2.0 license.
- ZAYA1-74B-Preview MoE language model preview with 4B active and 74B total parameters, demonstrating large-scale pretraining end-to-end on AMD with sliding-window attention and CCA. Pre-RL reasoning-base checkpoint trained on ~15T tokens across multiple phases. Apache 2.0 license.
- ZAYA1-VL-8B First Zyphra vision-language model. MoE architecture with 700M active / 8B total parameters, trained on ~140B vision-language tokens. Introduces vision-specific LoRA parameters and bidirectional attention for image tokens. Strong at document understanding, OCR, spatial perception, and GUI/computer-use tasks. Apache 2.0 license.
- ZAYA1-8B-Diffusion-Preview First MoE diffusion model converted from an autoregressive LLM (ZAYA1-8B-base), using the TiDAR recipe. Generates 16 tokens simultaneously, achieving up to 7.7x inference speedup. First diffusion-language model trained on AMD.
- ZONOS2 Next-generation real-time text-to-speech model with high-fidelity zero-shot voice cloning. Sparse MoE architecture with 900M active / 8B total parameters, trained on 6M+ hours of audio. Generates studio-quality 44.1 kHz audio via DAC tokens. Supports multilingual and code-switched generation. Apache 2.0 license.
- Zonos-v0.1 Beta text-to-speech model with high-fidelity voice cloning from 5-second audio samples. Released as 1.6B transformer and 1.6B hybrid variants under Apache 2.0; trained on ~200,000 hours of speech data.
- ZUNA 380M-parameter brain-computer interface foundation model for EEG data. Masked diffusion auto-encoder with 4D rotary positional encoding supporting arbitrary electrode layouts. Trained on ~2M channel-hours from 208 datasets. Foundation for noninvasive thought-to-text BCIs. Apache 2.0 license with MNE-compatible inference stack.
- Zamba2-7B State-of-the-art 7B hybrid SSM-Transformer small language model outperforming Mistral, Gemma, and Llama3 series in quality and performance at the 7B scale.
- Zamba2-mini 1.2B parameter state-of-the-art small language model with hybrid SSM-Transformer architecture. Fits in <700MB at 4-bit quantization for on-device AI deployments.
- Zamba2-Small (2.7B) 2.7B state-of-the-art small language model for on-device applications. Hybrid SSM-Transformer architecture achieving 2x speed and 27% reduced memory overhead.
- Zamba2-VL Family of open vision-language models built on the Zamba2 hybrid SSM-Transformer backbone, available at 1.2B, 2.7B, and 7B parameters. Competitive with leading Transformer-based VLMs at comparable scales while running substantially faster.
- ZR1-1.5B 1.5B parameter reasoning model trained extensively with reinforcement learning on coding and mathematics. Achieved parity with Claude 3 Opus on LCB-Generation while using 60% fewer tokens than R1-Distill-1.5B.
- Zyda-2 5-trillion-token high-quality open dataset for language modeling composed of filtered and cross-deduplicated DCLM, FineWeb-Edu, Zyda-1, and Dolma v1.7 Common Crawl portion. Processed using NVIDIA NeMo Curator (3 weeks reduced to 2 days). Powers the Zamba2 series.
- ZAYA1-base First large-scale Mixture-of-Experts foundation model trained entirely on an AMD platform (MI300X with Pensando Pollara networking), achieving over 750 PFLOPs of training performance. Pretrained on full-stack AMD compute, network, and system design.
Quantifiable outcome
- 2.6x throughput improvement (173M vs 66.30M tokens/sec) at 1,024 GPU scale and 128K context length with TSP vs matched TP+SP
- +6 more outcomes
Companies that use Zyphra
Customer profileSegments4 records
Ideal customer profiles4 records
Zyphra technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration6 records
AI capability14 records
Feature13 records
Zyphra partnerships and signals
Strategic signalScale indicators15 records
Recent moves7 records
Expansion highlights6 records
Zyphra competitors and assessment
Company assessmentDirect peers
- Mistral AI: European open-weight foundation model developer releasing competitive LLMs under Apache 2.0 with both open-source and enterprise cloud offerings. Closest direct comparable to Zyphra on open-weight strategy, multimodal model portfolio, and hybrid cloud + self-host distribution.
- DeepSeek: Chinese open-weight LLM lab producing frontier-tier MoE models (V3.2, R1) with strong reasoning and code benchmarks. Directly comparable on open-weight model releases and serves as both a peer and a hosted model on Zyphra Cloud, making it a uniquely bidirectional competitor.
- Together AI: Open-model-focused AI cloud platform offering serverless inference and dedicated GPU clusters for open-weight models. Closest analog to Zyphra's cloud + open-weight business model, though Together AI runs primarily on NVIDIA infrastructure rather than AMD.
- Cohere: Enterprise-focused foundation model provider offering proprietary LLMs and inference infrastructure to enterprises. Comparable on enterprise GTM and serving infrastructure, though Cohere's models are closed-source and primarily NVIDIA-based.
- AI21 Labs: Foundation model company developing both proprietary and open-weight LLMs (Jamba hybrid SSM-Transformer) targeting enterprise use cases. Notably shares Zyphra's hybrid architecture research lineage (Mamba/SSM) and comparable enterprise + open-source positioning.
- Anyscale: AI compute platform built on Ray for distributed training, inference, and serving at scale. Comparable infrastructure-layer offering to Zyphra Cloud for AI builders and enterprises needing scalable compute, though Anyscale is NVIDIA-only.
Broad incumbents
- OpenAI: Closed-source frontier AI lab with industry-leading model capabilities, enterprise distribution (ChatGPT Enterprise, API), and massive compute footprint. Sets the bar Zyphra must beat on capability and is the primary alternative for enterprise AI budgets.
- Anthropic: Frontier AI lab known for Claude models with strong enterprise and coding positioning. Represents the highest-end closed-source competitor in the enterprise inference market Zyphra Cloud targets.
- Meta AI (Llama): Creator of the Llama family of open-weight models that defined and still dominates the open-weight LLM category. Defines the ecosystem standard Zyphra must differentiate within, particularly via AMD-native positioning.
Emerging players
- ElevenLabs: Leading commercial text-to-speech and voice cloning platform. Most direct comparable to Zyphra's ZONOS2 product on capability, distribution, and enterprise TTS use cases, though ElevenLabs is closed-source and runs on NVIDIA infrastructure.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat6 records
Key risks6 records
Key highlights7 records
Customer concentration
Zyphra social profiles
Digital presenceZyphra financial estimates
Financial estimateRevenue estimate
Valuation estimate
Zyphra leadership team
Management profileNumber of profiles
Profiles7 records
Zyphra funding detail
Funding detailFunding overview
Funding rounds3 records
Investors1 record
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Zyphra M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Zyphra
What does Zyphra do?
Zyphra operates a full-stack open AI platform combining Zyphra Research (novel foundation model architectures including MoE with CCA attention and custom AMD-native training/inference kernels) and Zyphra Cloud (bare-metal AMD GPU infrastructure delivering serverless and dedicated inference for frontier open-weight models, the MAIA multiplayer agent platform, and custom hyperscale deployments). The company releases open-weight Apache 2.0 foundation models across language, vision-language, text-to-speech, diffusion, and BCI modalities, with proprietary parallelism and decoding schemes optimized for AMD MI300X/MI355X hardware.
Is Zyphra a public or private company?
Zyphra is a private company. It is classified as venture growth investor backed and is currently operating.
When was Zyphra founded?
Zyphra was founded in 2021. It employs 51 to 100 people.
Where is Zyphra based?
Zyphra is headquartered in Palo Alto, United States, in the North America region.
How does Zyphra make money?
Four revenue lines are on record. Zyphra Cloud Serverless Inference is the primary driver. The others are enterprise Dedicated Capacity and Custom Hyperscale Infrastructure, agent Platform and Compute Services and promotional Free Service (ZONOS2 launch).
Who are Zyphra's main competitors?
Direct peers on record are Mistral AI, DeepSeek, Together AI, Cohere, AI21 Labs and Anyscale. Broad incumbents are OpenAI, Anthropic and Meta AI (Llama). ElevenLabs is listed as an emerging player.
Does Zyphra have an API?
Yes. Public API available via Zyphra Cloud for hosted models including ZAYA1-8B and ZONOS2. The Terms of Use reference Zyphra APIs (used when accessing Zyphra Cloud, submitting prompts, or using other services). Models can also be self-hosted, with inference code published on GitHub for ZONOS2 (Python/SGLang-based) and an MNE-compatible inference stack for ZUNA. Access is also available through a hosted playground on Zyphra Cloud.
What industry is Zyphra in?
Zyphra's product category is AI Foundation Models & Cloud Platform. Its primary akta.pro industry code is HDAAACAB, Model Hosting, Serving & Inference Platforms, with a secondary code of HDAAAAAI, AI Compiler, Runtime & Kernel Optimization Software (CUDA/ROCm/XLA, graph compilers). Its NAICS code is 5182 and its SIC code is 7372.