Open-weight voice model Confucius4 survives screaming World Cup commentary stress test
dansuy_gaming · reddit · 2026-09-04
A voice-model enthusiast stress-tested the open-weight voice model Confucius4 with three brutal clips: a Spanish commentator screaming through a hat-trick call, an English commentator's held-breath-then-explosion moment, and a visibly shaken goalkeeper's post-match interview — all recently recorded World Cup commentary translated into other languages, with no scripts available.
These high-emotion clips typically expose cloned speech models (robotic screams or dead, context-free translation). Confucius4 takes the voice directly from audio rather than a transcript first. Results: short, high-emotion clips carried the trembling and emotion into the translation with little synthetic quality; longer sentences showed more synthetic artifacts.
More from Multimodal
- 1,300 photos rebuild Yokosuka port in 3DGS: LichtFeld Studio's new BG and exposure fixes put to test — janusch_patas · 2026-09-05
- GPT-6 Astra demoed building a dragon lair dungeon scene directly in Blender — majidmanzarpour · 2026-09-05
- Live AI-generated TV channel runs inside Minecraft, controlled via chat — isidentical · 2026-09-05
- ChatGPT photo fixes now look pro while keeping faces identical — 5 prompts shared — HeyAmit_ · 2026-09-05
- AI short film pits Street Fighter's Akuma against Baki's Yujiro — Ok-Vegetable-2455 · 2026-09-05
- AI video brings plush character Pillowdear to life on a balcony — AmyRoseFan_1234 · 2026-09-05