Small Parameter Voice Model Leads in Multilingual Capabilities
mattturck · x · 2026-07-16
A reshared post introduces GradiumAI's Phonon: despite having only 100 million parameters, it already leads in English. Now, its word error rates in French, German, and Spanish are also lower than NVIDIA's Magpie (357 million parameters) and Neuphonic's NeuTTS Nano (229 million parameters).
The post also emphasizes its support for high-quality voice cloning and its smaller model size. The original post includes comprehensive multilingual benchmarks and access links.
More from Multimodal
- NVIDIA ships Nemotron audio-native open weights in 2B and 30B sizes — victormustar · 2026-07-21
- HOMIE pairs Qwen3-VL-2B with Wan2.1 for human-object-centric video personalization — switch2stock · 2026-07-21
- Video Models Cut Ad Production Costs by 90-99%: Runway Enterprise Data — c_valenzuelab · 2026-07-21
- Pablo Stanley shares a full AI video workflow using ChatGPT, Gemini, Runway and CapCut — jdjohnson · 2026-07-21
- Meta AI text input now lets users interleave images with text — ezyang · 2026-07-21
- ShotPlan adds learnable planning tokens for cinematic multi-shot video generation — Tele-AI · 2026-07-21