Fish Audio says its voice clone now needs just 10 seconds of speech
aryanXmahajan · x · 2026-07-29
Fish Audio’s S2.1 Pro is being promoted as a developer-friendly voice cloning model that can copy a voice from just 10 seconds of speech.
- It preserves tone, rhythm, emotion, and accent, and supports 83 languages.
- The model offers real-time streaming and a developer API, with the post saying it is now free for developers.
- The author highlights use cases such as AI sales callers, support agents, multilingual podcasts, game characters, and real-time voice products.
Related event: Fish Audio Launches S2.1 Pro Voice Model with Ultra-Low Latency(4 posts)→
More from Multimodal
- Leaked timeline points to Hailuo 3, FLUX 3 and WAN 3 arriving within days — koltregaskes · 2026-07-29
- A local iPhone photo editor enters beta with on-device vision models — measure_plan · 2026-07-29
- Gemini Flash turns a distressed robot photo into a meme — Bishopkilljoy · 2026-07-29
- Open-source app uses Gemini Video Understanding to pull highlight reels automatically — icnahom · 2026-07-29
- Researchers want agent runs to end with short explainer videos, not text walls — airesearch12 · 2026-07-29
- WAN Bernini plus Prompt Relay gives finer control over 10–15 second videos — Sudden_List_2693 · 2026-07-29