WhissleAI
Whissle Inc. operates a voice AI platform anchored on its proprietary META-1 model, streaming transcription with inline emotion, intent, and demographic metadata for contact centers, sales coaching, and code-mixed language deployments via self-hosted Docker and macOS applications.
- Company typePrivate
- Founded2023
- HeadquartersLos Angeles, United States
- Headcount1–10
- GTM typeB2B
- OfferingSoftware
What WhissleAI does
Whissle Inc. (operating as WhissleAI) is a privately held US-incorporated AI company that has built a real-time multi-modal speech and voice intelligence platform founded on its proprietary META-1 discriminative model. META-1 emits transcription tokens interleaved with action tokens for emotion (7 classes), intent (20+ categories), behavior (26 types), age range, gender, speaker change, speech rate, filler words, and named entities inside a single CTC forward pass, then pairs the output with KenLM beam-search decoding and an LLM layer (configurable across Claude, Gemini, Groq, or local llama.cpp/vLLM/Ollama). The shipping stack includes Whissle Gateway — a self-hostable Docker container that bundles ASR, TTS via Kokoro 82M, speaker diarization, and a dual-lane video intelligence layer (MediaPipe fast lane plus semantic vision LLM) on a single GPU — and a macOS AI browser (Whissle Browser) embedding an in-browser Lulu voice companion, with Agents Studio (visual agent builder with voice cloning) and a Whissle Cloud API both slated to return or launch.
The business model is API-first and product-led with two monetization streams: usage-based cloud pricing at $0.0043 per audio minute that bundles all metadata rather than selling emotion/intent/demographics as add-ons, and perpetual self-hosted gateway licenses that eliminate per-minute fees for high-volume enterprise customers. Distribution is developer-first — open-source footprint across 40+ GitHub repositories, 55+ HuggingFace models and datasets, a community Slack under the affiliated Whissle.org 501(c)(3), a one-line Docker install script, and peer-reviewed blog benchmarks — paired with a solutions page targeting contact centers and collections (including code-mixed Hindi-English), sales coaching, behavioral analytics, and 3D avatar agents. The proprietary 501(c)(3) spin-out (Whissle.org) and Google Summer of Code cohort indicate a deliberate research-and-community strategy layered onto the commercial entity.
Material commercial signals are thin: no revenue, ARR, headcount, funding round, or named customer is disclosed in the available source material, and Whissle Cloud API and Whissle Web App are explicitly marked "temporarily down" while on-prem capabilities are reinforced. Peer-reviewed publications at EMNLP 2023 (Industry) and EMNLP 2025, plus the arXiv 1SPU 2024 release and benchmark wins against Deepgram Nova-3 and AssemblyAI, provide external validation of the underlying research, but the absence of disclosed financials, hiring cadence, and enterprise logos means the current scale and growth pace cannot be precisely characterized from the input.
WhissleAI firmographics
Firmographics- Name
- WhissleAI
- Legal name
- Whissle Inc.
- Website
- https://whissle.ai
- Company type
- Private
- Founded year
- 2023
- Operating status
- Operating
- Headcount range
- 1–10 employees
- Short description
- Whissle Inc. operates a voice AI platform anchored on its proprietary META-1 model, streaming transcription with inline emotion, intent, and demographic metadata for contact centers, sales coaching, and code-mixed language deployments via self-hosted Docker and macOS applications.
- Ownership category
- akta.pro rank
WhissleAI industry classification
Industry- Product category
- Voice AI / Speech Recognition Platform
- NAICS
- Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (5182)
- SIC
- Services-Prepackaged Software (7372), Services-Computer Programming, Data Processing, Etc. (7370)
- akta.pro primary industry
- Conversational AI Platforms (chat/voice bots, orchestration) (HDAAAFAC)
- akta.pro secondary industries
- Model Hosting, Serving & Inference Platforms (HDAAACAB), On-Device Speech & Audio AI (wake word, ASR, enhancement) (HDAAAJAF)
Keywords
Where WhissleAI is headquartered
LocationHeadquarters
- HQ city
- Los Angeles
- HQ country
- United States
- HQ region
- North America
Markets served
WhissleAI business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Technology or R&D, Infrastructure, Personnel, Marketing or Sales
Revenue model
- Cloud API Usage: Usage-based pricing at $0.0043 per audio minute for cloud API streaming transcription with full metadata (emotion, intent, demographics, speech rate).
- Self-Hosted Gateway License: Self-hosted Docker deployment eliminates per-minute costs. Organizations deploy on their own GPU infrastructure with no recurring fees to Whissle.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Usage-based | Pay-as-you-go | Cloud API Streaming |
| One time/ perpetual license | Multi-year contract | Self-Hosted Gateway |
Go-to-market motion2 records
Distribution channels4 records
Marketing channels6 records
WhissleAI product offering
Product offeringCore offering
WhissleAI builds a voice AI platform offered both as a usage-based cloud API and as a self-hostable Docker deployment, anchored by its proprietary META-1 multi-modal discriminative foundation model. It also ships consumer-facing surfaces — a macOS AI browser, a macOS app, and a web app — that expose the same underlying speech and agent capabilities to end users.
Product overview
WhissleAI offers a modular instant intelligence platform that bridges discriminative and generative AI — converting any stream (audio, text, or video) into structured transcripts, emotion, intent, and actionable insights in real time. The core offering centers on Whissle Gateway (self-hosted Docker container running full-stack voice AI: ASR, LLM, TTS via Kokoro, speaker diarization, and video intelligence) and Whissle Browser (macOS AI browser with ambient voice companion Lulu). The META-1 foundation model is the multi-modal architecture underlying the platform, extracting metadata tokens inline during CTC decoding in a single pass. Whissle Cloud API (returning soon) and Whissle Web App provide cloud-hosted access. Agents Studio (coming soon) adds a visual agent builder with voice cloning and telephony integration. All products share the Stream2Action architecture, converting any input stream into structured JSON with auto-dispatch, human escalation, and webhook actions.
Differentiator
Problem solved
Functional benefit
Brands
- Whissle Browser: AI browser with ambient voice intelligence, adaptive theming, inline annotations, and Lulu built in.
- Whissle Gateway
- Whissle macOS App
- Agents Studio
- Lulu
- META-1
Quantifiable outcome
- 30-second average call resolution vs 2+ minutes with traditional IVR
- +5 more outcomes
Companies that use WhissleAI
Customer profileSegments4 records
Ideal customer profiles3 records
WhissleAI technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration5 records
AI capability12 records
Feature6 records
WhissleAI partnerships and signals
Strategic signalPartnerships
Five partnerships are on record, tiered core.
- Red Hen LabcoreCollaboration on open-source research through Google Summer of Code 2024. Three students contributed to multilingual language models, RAG pipelines, and visual-aware speech recognition. Joint publication on The Red Hen anonymizer and de-identifying audiovisual recordings.
- Google Summer of CodecoreGSoC 2024 program supporting student contributions to open-source AI research. Whissle participated with three students working on multilingual news LLM fine-tuning, RAG pipeline development, and visual-aware speech recognition in noisy settings.
- PipecatcorePipecat powers the voice calling platform inside Whissle Gateway. Handles WebRTC browser voice chat and integrates with Twilio for phone calls. Pipecat service runs on port 8000 inside the container.
- Kokoro TTScoreKokoro 82M text-to-speech engine embedded in Whissle Gateway. Non-autoregressive, single forward pass, 55 voices, 10 languages. Sub-200ms TTFB on CPU.
- Anthropic ClaudecoreClaude API integration for LLM processing in voice pipeline. Configurable LLM provider (claude, gemini, groq, or local llama.cpp/vLLM/Ollama). Default model: claude-sonnet-4-6.
Scale indicators7 records
Recent moves7 records
Expansion highlights6 records
WhissleAI competitors and assessment
Company assessmentEmerging players
- ElevenLabs: ElevenLabs leads in neural text-to-speech and conversational AI agents with emotion/voice cloning capabilities. Adjacent to Whissle on the conversational voice-agent stack (TTS + agent loop) but optimization-led approach differs from Whissle's ASR metadata-first architecture.
- Vapi: Vapi is a developer platform for building and deploying voice AI agents with telephony integration, comparable to Whissle's planned Agents Studio. Competes on agent orchestration and time-to-deploy for voice-agent startups rather than under-the-hood ASR/model innovation.
Direct peers
- Deepgram: Deepgram is the closest direct competitor — cloud speech-to-text API with per-minute usage pricing, metadata features (sentiment, intent, diarization), and self-hosted/on-prem deployment options. Whissle explicitly matches Deepgram's $0.0043/min base rate while bundling metadata that Deepgram charges as add-ons.
- Hume AI: Hume AI builds voice AI focused on emotional understanding, expression measurement, and empathic voice interfaces — directly overlapping Whissle's emotion/intent metadata thesis but as a generative affective computing layer over LLM speech rather than an ASR-first stack.
- Speechmatics: Speechmatics offers cloud and on-premise ASR with strong multilingual coverage including code-switching and accent robustness. Competes head-to-head with Whissle on accented/code-mixed speech accuracy and on-prem deployment for regulated industries.
- Rev.ai: Rev.ai provides async and streaming speech-to-text APIs with human-transcription lineage. Targets the same enterprise transcription and call-analytics use cases as Whissle with similar developer-first GTM.
- AssemblyAI: AssemblyAI is a unified AI speech API offering ASR, sentiment analysis, content moderation, and LLM-based summarization over voice. Targets the same developer and enterprise call-center buyers with similar per-minute usage pricing and metadata features that Whissle benchmarks against (25.6% failure rate cited).
Broad incumbents
- OpenAI (Whisper / Realtime API): OpenAI's Whisper (batch) and GPT-4o Realtime (streaming multimodal voice) serve as a well-funded default option for transcription and voice agents. Competes on quality and ecosystem ubiquity rather than specialized metadata extraction or on-prem deployment.
- Google Cloud Speech-to-Text: Google Cloud STT offers streaming and batch ASR with broad language coverage and native Vertex AI integration. A broad incumbent competing on enterprise procurement convenience and bundled AI tooling rather than metadata-first design.
- Amazon Transcribe: AWS Transcribe provides streaming and batch speech recognition with Call Analytics (sentiment, categories, PII redaction) inside the AWS ecosystem. Competes on cloud-bundle lock-in for enterprises already on AWS rather than self-host metadata differentiation.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat5 records
Key risks6 records
Key highlights7 records
Customer concentration
WhissleAI social profiles
Digital presenceWhissleAI financial estimates
Financial estimateRevenue estimate
Valuation estimate
WhissleAI leadership team
Management profileNumber of profiles
WhissleAI subsidiaries and ownership
Company hierarchySubsidiaries1 record
WhissleAI funding detail
Funding detailFunding overview
Funding rounds
Investors
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
WhissleAI M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about WhissleAI
What does WhissleAI do?
WhissleAI builds a voice AI platform offered both as a usage-based cloud API and as a self-hostable Docker deployment, anchored by its proprietary META-1 multi-modal discriminative foundation model. It also ships consumer-facing surfaces — a macOS AI browser, a macOS app, and a web app — that expose the same underlying speech and agent capabilities to end users.
Is WhissleAI a public or private company?
WhissleAI is a private company. It is classified as founder individual operated bootstrapped and is currently operating.
When was WhissleAI founded?
WhissleAI was founded in 2023. It employs 1 to 10 people.
Where is WhissleAI based?
WhissleAI is headquartered in Los Angeles, United States, in the North America region.
How does WhissleAI make money?
Two revenue lines are on record. Cloud API Usage is the primary driver. The others are self-Hosted Gateway License.
Who are WhissleAI's main competitors?
Emerging players on record are ElevenLabs and Vapi. Direct peers are Deepgram, Hume AI, Speechmatics, Rev.ai and AssemblyAI. Broad incumbents are OpenAI (Whisper / Realtime API), Google Cloud Speech-to-Text and Amazon Transcribe.
Does WhissleAI have an API?
Yes. Whissle Gateway runs seven services in a single Docker container: PostgreSQL (port 5432), ASR (port 8001, REST + WebSocket), Video Intelligence (port 8002, REST + WebSocket), TTS/Kokoro (port 8003, WebSocket), Pipecat voice calling (port 8000, REST + WebSocket), Agent/LLM processing (port 8765, REST + SSE), and Gateway proxy (port 9000, REST + WebSocket). Authentication uses built-in API tokens (auto-generated admin token, user-level JWT cookies for voice calling). WebSocket streaming endpoints support real-time ASR with metadata (emotion, intent, speaker change, demographics, speech rate, filler words), WebRTC voice calling, video intelligence (face emotion, gaze, gestures, scene), and TTS. REST endpoints cover batch transcription, diarization, video batch analysis, agent chat, auth, organizations, invitations, and database. Supported LLM providers: Claude (Anthropic), Gemini (Google), Groq, or local (llama.cpp / vLLM / Ollama). Configuration via environment variables with options for HIPAA mode, Twilio, Google OAuth, and multi-tenancy. Developer documentation is at www.whissle.ai/gateway/docs.
What industry is WhissleAI in?
WhissleAI's product category is Voice AI / Speech Recognition Platform. Its primary akta.pro industry code is HDAAAFAC, Conversational AI Platforms (chat/voice bots, orchestration), with a secondary code of HDAAACAB, Model Hosting, Serving & Inference Platforms. Its NAICS code is 5182 and its SIC code is 7372.