HC
job looker
Small-Business (50 or fewer emp.)
"Fish Audio: Expressive, Low-Latency TTS with Open Weights and a Developer-Friendly API"
4/5
What do you like best about Fish Audio?

What I like most about Fish Audio is how it balances a clean, no-frills web UI (playground, Story Studio for multi-character long-form narration, and a 2M+ community voice library) with genuinely production-grade engineering: the S2/S2.1-Pro models deliver expressive, emotion-tagged TTS and 10–15s zero-shot voice cloning that beats ElevenLabs on prosody in several non-English languages, while the REST + WebSocket API with official Python/TypeScript SDKs, LiveKit/Pipecat hooks, llms.txt/OpenAPI specs and an agent skill for Cursor/Claude/Codex make it one of the most developer-friendly voice stacks out there, all running at sub-500ms streaming latency suitable for real-time voice agents and conversational AI. Performance-wise it scales from free prototyping to pay-as-you-go API at ~$15 per 1M UTF-8 bytes and Plus/Pro tiers from ~$11–$100/month, which is roughly a sixth of ElevenLabs’ cost for comparable output, and the open-weight Fish-Speech/S1/S2 models let teams self-host in VPC or air-gapped environments when they need data residency. After-sales support is lighter than enterprise incumbents—email plus docs, no white-glove SLA on lower tiers—but the open-source heritage (So-VITS-SVC, Bert-VITS2 lineage), transparent blind-test benchmarks (Audio Turing score 0.515 on S2.1 Pro) and active Discord/Hacker News/Reddit community partly compensate for the lack of hand-holding. On agents specifically, it’s a standout: inline direction tags like [whisper] or [chuckle] travel inside the text, ASR+TTS+clone live in one stack, and interruption-aware streaming makes it easy to wire into Retell/HeyGen-style assistants, so the thing I’d keep over any competitor is the combination of low-cost expressive quality, open weights, and an API surface that treats voice as a programmable performance layer rather than a black-box narration button. Review collected by and hosted on G2.com.

What do you dislike about Fish Audio?

What I like least about Fish Audio is that the things that make it appealing to developers—open weights, byte-metered API, inline emotion tags, concurrency-gated tiers—are exactly what make it feel rough around the edges for everyone else: the web UI is functional but unpolished compared to ElevenLabs’ Studio, with no real waveform editor, light post-processing, or non-technical project-management layer, so creators still bounce to external DAWs; the 2M+ community voice library is huge but unmoderated, so commercial users have to manually audition voices and carry the licensing/consent risk themselves since public voices are not automatically cleared; integrations cover REST/WebSocket/Python/TS plus LiveKit/Pipecat/MCP hooks but stop short of first-class LangChain/LlamaIndex wrappers, deeper no-code connectors, or the bundled STT-LLM-TTS orchestration that more mature voice-agent stacks ship, leaving you to wire interrupt handling, retry, and concurrency math yourself against a prepaid-spend concurrency ladder rather than simple QPS; performance is genuinely fast (sub-150ms TTFA) yet English long-form stability and smaller-language fidelity still trail ElevenLabs, emotion tags occasionally fire inconsistently, and 15-second clones degrade sharply on noisy reference audio; pricing is cheap on paper (~$15/M UTF-8 bytes) but the credit/byte/audio-hour mixing, non-rolling monthly quotas, and free tier capped at a few thousand characters per day make real cost modeling annoying for irregular teams; after-sales is email-plus-docs with no published SLA on self-serve tiers, and the compliance story (no clear SOC 2/HIPAA/ISO posture, thin DPA, ambiguous prompt-training opt-out) blocks healthcare/finance/government buyers despite the self-host escape hatch; and on agents specifically, while inline tags and streaming latency are great for prototyping, production voice agents still need you to build turn-taking, barge-in, session memory, and fallback logic on top, with S2’s removed LoRA fine-tune path and tag-placement sensitivity meaning you trade ElevenLabs-style “set stability/clarity/style sliders and forget it” for more prompt engineering and more eval work—so the honest downside is that Fish Audio is a fantastic programmable voice engine, not a finished enterprise narration or conversational-platform product. Review collected by and hosted on G2.com.

See what 116 reviewers think of Fish Audio

4.5 out of 5 · Verified reviews from real users

Read all reviews