Cartesia's Sonic 3.6 Clones a Voice From 15 Seconds of Audio, Speaks 44 Languages
omarsar0 · x · 2026-10-03
Elvis Saravia tested Cartesia's latest voice model Sonic 3.6: a voice clone was ready in seconds from a 15-second clip, then spoke fluent Japanese — a language he doesn't speak. The model supports 44 languages, and he suggests using it to make content accessible across languages.
More from Multimodal
- Fable 5.5 generates a stunning 40,000-year art history animation — dotey · 2026-10-04
- Looped-DiT: 260M looped model beats 6.5x larger ones with 4.9x less inference compute — burny_tech · 2026-10-04
- Augie's AI Image Browser indexes ComfyUI metadata and offers leak-free exports for selling — spanktastic0x · 2026-10-04
- Music cover generated on 8GB VRAM: YuE2 + LTX-2 render in 17 minutes — big-boss_97 · 2026-10-04
- Image-to-image character translation on Runway turns host into an 'Alternative Late Show' — c_valenzuelab · 2026-10-04
- Immunologist generates 2-minute immunology history video entirely in code with Claude — DeryaTR_ · 2026-10-04