Fish Audio S2
Fish Audio S2, operated by Hanabi AI Inc., is an AI voice generation platform offering expressive text-to-speech, voice cloning, and speech-to-text across 30+ languages. It serves content creators, developers, and enterprises through a freemium website, REST API, and compliance-ready enterprise contracts.
- Company typePrivate
- Founded2025
- HeadquartersMountain View, United States
- Headcount11–50
- GTM typeB2B and B2C
- OfferingSoftware
What Fish Audio S2 does
Fish Audio S2, operated by Hanabi AI Inc. (a Delaware-incorporated company headquartered in Mountain View, California), is an AI voice generation platform that provides text-to-speech, speech-to-text, voice cloning, and voice design through a self-serve website (fish.audio), a REST API, and mobile applications. The company maintains a proprietary model family — Fish Speech, Fish Audio S1, Fish Audio S2, and the flagship Fish Audio S2.1 Pro — that supports real-time, emotionally controllable voice generation across more than 30 languages, with native-quality pronunciation in 8 (English, Japanese, Korean, Chinese, French, German, Arabic, Spanish).
The product surface is built around a core TTS engine with Voice Cloning that creates a digital voice model from 10-15 seconds of audio, Voice Design for prompt-based voice creation, Story Studio for ACX/Audible-spec-compliant audiobook production, and a Voice Library marketplace containing over 2,000,000 community-uploaded voices. Supporting modules include Voice Changer, Audio Separation, Audio Translation, Sound Effects, Speech-to-Text transcription, and a Voice Agent enterprise solution for customer-support applications. Development is sustained by an open-source community through GitHub (fishaudio), Discord, and the OpenAudio research arm, which together with 18+ integration partnerships (HeyGen, OpenArt, Novita AI, Viggle, VoiceDrop, Pictoria, and others) form a distribution ecosystem.
Fish Audio monetizes through a freemium tiered subscription (Free, Plus at $11/month, Pro at $75/month, Max at $749/month), pay-as-you-go API pricing ($15 per million UTF-8 bytes for TTS, $0.36 per audio hour for ASR), and custom enterprise contracts that bundle SOC 2 Type II compliance, zero data retention, on-premise deployment, and custom SSO. Its go-to-market is primarily product-led and self-serve, with developer-focused API documentation for technical buyers and direct enterprise sales for compliance-driven customers. The platform serves three primary segments — content creators and YouTubers, developers building voice-enabled applications, and enterprises with regulatory requirements — with gaming, animation, and audiobook publishing as secondary verticals.
Fish Audio S2 firmographics
Firmographics- Name
- Fish Audio S2
- Legal name
- Hanabi AI Inc.
- Website
- https://fishaudio.xyz
- Company type
- Private
- Founded year
- 2025
- Operating status
- Operating
- Headcount range
- 11–50 employees
- Short description
- Fish Audio S2, operated by Hanabi AI Inc., is an AI voice generation platform offering expressive text-to-speech, voice cloning, and speech-to-text across 30+ languages. It serves content creators, developers, and enterprises through a freemium website, REST API, and compliance-ready enterprise contracts.
- Ownership category
- akta.pro rank
Fish Audio S2 industry classification
Industry- Product category
- AI Voice Synthesis Software
- NAICS
- Sound Recording Industries (5122), Sound Recording Studios (51224), Other Sound Recording Industries (512290), Record Production and Distribution (51225)
- akta.pro primary industry
- Voice Cloning, Dubbing & Audio Generative Media (HDAAAFAK)
- akta.pro secondary industry
- Voiceover Recording & Production (Narration, Promo, Commercial VO) (MPAHAEAK)
Keywords
Where Fish Audio S2 is headquartered
LocationHeadquarters
- HQ city
- Mountain View
- HQ country
- United States
- HQ region
- North America
Offices1 record
Markets served
Fish Audio S2 business model
Business model- GTM type
- B2B and B2C
- Offering type
- Software
- Cost components
- Technology or R&D, Infrastructure, Personnel, Operations, Marketing or Sales
Revenue model
- Subscription Plans: Tiered subscription model with Free, Plus ($11/mo), Pro ($75/mo), and Max ($749/mo) tiers providing monthly credit allocations for voice generation. Annual billing offers 33% savings. Credits expire monthly and do not roll over.
- API Usage-Based Pricing: Pay-as-you-go pricing for API access based on actual usage. TTS models priced at $15.00 per million UTF-8 bytes, Speech Recognition at $0.36 per audio hour, Voice Design at $0.01 per successful request. No subscription fees or monthly minimums.
- Enterprise Services: Custom volume pricing for organizations with compliance needs including Zero Data Retention, On-Premise Deployment, SOC2 Compliance, Custom SSO, and organization-level controls.
- Freemium to Paid Conversion: Free tier provides limited generations monthly for personal use; paid plans required for commercial use and monetization of generated content.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Freemium | Monthly | Free tier with limited generation for personal use |
| Subscription | Annual | For creators and professionals with increased generation limits |
| Subscription | Annual | For power users and businesses with high-volume needs |
| Subscription | Annual | For teams with large-scale production needs |
| Subscription | Annual | Custom enterprise pricing with compliance features |
| Usage-based | Pay-as-you-go | API pay-as-you-go pricing for developers |
Go-to-market motion2 records
Distribution channels4 records
Marketing channels6 records
Fish Audio S2 product offering
Product offeringCore offering
Fish Audio S2 is an AI voice platform that converts text into expressive, multilingual speech and clones voices from short audio samples. It sells subscriptions and a usage-based API to individual creators, developers, and enterprises, complemented by purpose-built tools such as Story Studio for audiobooks and a Voice Library of over 2,000,000 community-uploaded voices.
Product overview
Fish Audio S2 is a unified AI voice generation platform organized around a core text-to-speech engine (Fish Audio S2) with multiple integrated modules and tools. The platform's core products include Text-to-Speech (the primary TTS engine with emotion control and multilingual support) and Speech-to-Speech (ASR transcription). These are complemented by voice cloning tools that can create digital voice models from 10-15 seconds of audio, a Story Studio for audiobook production, Voice Changer for real-time modification, Audio Separation and Translation services, and a Sound Effects library. The platform also includes a Voice Library marketplace with over 2,000,000 community voices and a developer API for programmatic access. Enterprise solutions include Voice Agent for customer support applications. The platform supports 8+ languages and offers tiered pricing from free (8,000 credits/month) to Max ($749/month for 25M credits).
Differentiator
Problem solved
Functional benefit
Brands
- Fish Audio S2: The latest generation voice model with enhanced emotional control and expressiveness.
- Fish Audio S1
- OpenAudio
Products and services
- Text-to-Speech AI text-to-speech engine that converts text into natural, broadcast-quality voice audio with emotion control and multilingual support across 30+ languages; used by creators, developers, and enterprise teams for video voiceovers, audiobooks, and conversational agents.
- Speech-to-Text Automatic speech recognition service that transcribes audio to text with support for multispeaker audio, emotion tags, and natural language descriptions; intended for developers and content teams needing transcription of voice input.
- Voice Cloning Standalone voice cloning product that creates a digital model of a person's voice from 10–15 seconds of audio and generates speech in that voice across multiple languages; targeted at creators, game studios, and brand teams building signature voices.
- Story Studio Audiobook production tool that generates publish-ready long-form audio with chapter-level control, emotion tags, and pacing; aimed at indie authors and audiobook publishers who need ACX/Audible-compliant output without a recording booth.
- Fish Audio API REST API offering programmatic access to text-to-speech, speech-to-text, voice cloning, and voice design models with pay-as-you-go pricing ($15.00 per million UTF-8 bytes for TTS, $0.36 per audio hour for ASR, $0.01 per voice design request) and tiered concurrent-request limits; used by developers integrating voice AI into applications.
- Voice Agent (Enterprise) Enterprise voice agent solution for customer support and virtual agents that combines ultra-low-latency synthesis with tone and emotion control; aimed at organizations needing compliant, branded conversational voice experiences.
Quantifiable outcome
- Cost reduction of 90-95% compared to traditional voice actors
- +2 more outcomes
Companies that use Fish Audio S2
Customer profileNamed customers4 records
Segments5 records
Ideal customer profiles4 records
Fish Audio S2 technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
AI capability4 records
Feature8 records
Fish Audio S2 partnerships and signals
Strategic signalPartnerships
One partnership is on record.
- StripecorePayment processing partner for subscription billing, payment processing, and financial transactions. Handles credit card information and payment processing for Fish Audio services.
Scale indicators5 records
Recent moves6 records
Expansion highlights7 records
Fish Audio S2 competitors and assessment
Company assessmentDirect peers
- ElevenLabs: Direct competitor offering AI text-to-speech, voice cloning, and multilingual dubbing. Fish Audio maintains explicit comparison pages against ElevenLabs across features and pricing, positioning on expressiveness, language coverage, and cost.
- Resemble AI: Specialist voice-cloning and real-time TTS platform targeting enterprise and developer use cases. Comparable on voice cloning, custom voices, and API access, with overlapping enterprise compliance positioning.
- Inworld AI: AI character voice and dialogue platform for games, animation, and interactive experiences. Overlaps Fish Audio's Character Voices and Gaming & Animation segments on expressive, emotion-rich TTS.
- Cartesia AI: Real-time generative voice AI with low-latency streaming TTS targeting developers and voice-agent builders. Fish Audio maintains a public comparison page and competes on latency, expressiveness, and API ergonomics.
- Play.ht: AI voice generator with large voice libraries, voice cloning, and multilingual TTS for content creators and publishers. Directly overlaps Fish Audio's creator + developer segments and Story Studio-style audiobook workflows.
- Hume AI: Conversational AI platform with emotionally aware voice (EVI) targeting voice agents and customer experience. Comparable on emotion-controlled TTS and the voice-agent enterprise use case Fish Audio explicitly markets.
- Murf AI: Enterprise-focused AI voice platform for video voiceovers, e-learning, and corporate narration. Comparable on emotion/tone control, voice library, and tiered subscription targeting business content production.
- Speechify: AI text-to-speech platform with consumer and enterprise tiers, voice cloning, and creator tools. Direct overlap on narration, audiobook, and video voiceover use cases that Fish Audio's Story Studio and TTS target.
Broad incumbents
- OpenAI: OpenAI ships TTS and Voice Engine capabilities bundled into its broader AI platform. Fish Audio differentiates on emotion control, voice cloning depth, and per-seat pricing for creators and developers outside the OpenAI ecosystem.
- Google Cloud Text-to-Speech: Hyperscaler TTS service with broad language coverage and enterprise SLAs. Represents the bundled, well-capitalized incumbent that Fish Audio's API pricing, emotion control, and voice cloning differentiation must outcompete on use-case depth.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat5 records
Key risks6 records
Key highlights6 records
Customer concentration
Fish Audio S2 social profiles
Digital presenceFish Audio S2 compliance and trust
Trust signalCompliance1 record
Fish Audio S2 financial estimates
Financial estimateRevenue estimate
Valuation estimate
Fish Audio S2 leadership team
Management profileNumber of profiles
Fish Audio S2 funding detail
Funding detailFunding overview
Funding rounds1 record
Investors11 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Fish Audio S2 M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Fish Audio S2
What does Fish Audio S2 do?
Fish Audio S2 is an AI voice platform that converts text into expressive, multilingual speech and clones voices from short audio samples. It sells subscriptions and a usage-based API to individual creators, developers, and enterprises, complemented by purpose-built tools such as Story Studio for audiobooks and a Voice Library of over 2,000,000 community-uploaded voices.
Is Fish Audio S2 a public or private company?
Fish Audio S2 is a private company. It is classified as founder individual operated bootstrapped and is currently operating.
When was Fish Audio S2 founded?
Fish Audio S2 was founded in 2025. It employs 11 to 50 people.
Where is Fish Audio S2 based?
Fish Audio S2 is headquartered in Mountain View, United States, in the North America region.
How does Fish Audio S2 make money?
Four revenue lines are on record. Subscription Plans are the primary driver. The others are API Usage-Based Pricing, enterprise Services and freemium to Paid Conversion.
Who are Fish Audio S2's main competitors?
Direct peers on record are ElevenLabs, Resemble AI, Inworld AI, Cartesia AI, Play.ht, Hume AI, Murf AI and Speechify. Broad incumbents are OpenAI and Google Cloud Text-to-Speech.
Does Fish Audio S2 have an API?
Yes. Fish Audio offers a REST API for text-to-speech, speech-to-text, voice cloning, and voice design. Pay-as-you-go pricing with no subscription fees or monthly minimums. TTS models include s2.1-pro ($15.00/M UTF-8 bytes), s2.1-pro-free ($0.00), s2-pro ($15.00/M), s1 ($15.00/M). ASR (transcribe-1) costs $0.36/audio hour. Voice Design (voice-design-1) costs $0.01 per successful request. Concurrent request limits: Starter (<$100 paid): 5 requests, Elevated (≥$100): 15 requests, High Volume (≥$1,000): 50 requests, Enterprise: custom limits. Developer documentation is at docs.fish.audio/api-reference/introduction.
What industry is Fish Audio S2 in?
Fish Audio S2's product category is AI Voice Synthesis Software. Its primary akta.pro industry code is HDAAAFAK, Voice Cloning, Dubbing & Audio Generative Media, with a secondary code of MPAHAEAK, Voiceover Recording & Production (Narration, Promo, Commercial VO). Its NAICS code is 5122.