Fish Audio says its new voice model clones a speaker from 5 seconds of audio
JafarNajafov · x · 2026-07-29
- Fish Audio says its new S2.1 Pro can clone a voice from 5 seconds of audio.
- The company claims it is 2× faster than Cartesia and costs 1/6 of ElevenLabs.
- It also positions the model around expressive speech control, including emotion, intonation, and pacing.
- The launch comes with a disclosed $52M seed round.
- Fish Audio says customers such as HeyGen, LiveKit, Retell, Sanas, and OpenArt already run its model in production.
Related event: Fish Audio Launches S2.1 Pro Voice Model with Ultra-Low Latency(4 posts)→
More from Venture
- Will Brown reportedly signs a $200 compute deal with OpenAI — willccbb · 2026-07-29
- Agentic AI is becoming a revenue-defense strategy, not just an ROI tool — firstadopter · 2026-07-29
- Chip stocks slide as investors cool on the AI trade — JumpCrisscross · 2026-07-29
- China’s venture market runs on “equity in name, debt in substance” and 2%–5% FA fees — deedydas · 2026-07-29
- An AI bubble debate turns on sample size, not just valuations, says a French press column — emmanuelvivier · 2026-07-29
- Bittensor’s $250M annual emissions are a tiny fraction of Big Tech’s AI spend — bittingthembits · 2026-07-29