Fish Audio Makes S2.1 Pro Voice Model Free, Costs 1/6 of ElevenLabs

rohanpaul_ai · x · 2026-07-30

Fish Audio has made its S2.1 Pro voice cloning model free for a month. The model requires only 10-15 seconds of audio to clone a voice, supports 83 languages, and features a low latency of 90ms with word-level pronunciation and pause control.

Architecture & Ecosystem: It uses a Dual-AR design (a 4B-parameter model for semantics/emotion and a 400M-parameter model for acoustic details) balancing speed and expressiveness. Built on their open-source project Fish Speech, they continue to offer open-weight models for self-hosting. Pricing is roughly 1/6th of ElevenLabs.

Core Advantage: Tuned for real conversation, the model is designed to survive interruptions, corrections, laughter, and language switches.

Related event: Fish Audio Launches S2.1 Pro Voice Model and Raises $52M(11 posts)→

Original post →

More from Models

Models channel →