Sopro V2 Turbo 2610: cleaner voice cloning, 120M model, ~300ms first audio on CPU
SammyDaBeast · reddit · 2026-10-04
Open-source voice cloning model Sopro V2 Turbo ships the 2610 interim update, mainly fixing roughness and break-up on cloned voices.
Key points:
- Same 120M model, same speed: 300ms to first audio on a laptop CPU, with true streaming
- Apache-2.0; supports English, European Portuguese, French, German; more languages and voice controls planned
- Known weaknesses: very high-pitched/cartoon voices, noisy reference audio, some OOD voices; author soliciting failed samples
- One-line local run via uvx, plus repo, HuggingFace weights, in-browser demo, and an evals blog post
A fit for those who like F5-TTS but want a much lighter model that streams and runs comfortably on CPU.
More from Multimodal
- AI talk show video: Uncle Sam roasts a $20,000 coffee machine — UnfilteredTalkOffice · 2026-10-04
- Creator shares Seedance 2.5 prompt for surreal mirror-to-ocean video — azed_ai · 2026-10-04
- AI-generated creature horror short film: The Ants — Attacks on Man — vishalstudiosverse · 2026-10-04
- One-Word Prompt Swap: Midjourney Renders a 'Glass Shepherd' Leading a Flock Through Fog — tisch_eins · 2026-10-04
- No video editor needed: dev turns app UI docs into a 30s motion graphics video with Codex + GPT-6.1 Sol — No-Hunter9792 · 2026-10-04
- Kling 4.0 Flash video model launched, teasing faster generation — azed_ai · 2026-10-04