Moshi’s full-duplex voice model predates today’s hottest speech AI buzzword
mattturck · x · 2026-07-25
- A discussion around full-duplex voice AI argues that the term is now hot, even though Moshi had already been built as an end-to-end conversational model before the label became common.
- Neil Zegh remarks that the team originally set out to build a conversational system, without prior dialogue-system experience, and only later realized their work fit a broader technical concept.
- The linked talk frames Moshi as an early speech-to-speech, full-duplex assistant, compares it with newer live speech systems, and asks why this design has not become ubiquitous yet.
More from Multimodal
- Midjourney V8.2 adds personalization and shows off stylized image outputs — Mr_AllenT · 2026-07-27
- Midjourney’s image variety draws a Krea 2 comparison and asks how to reproduce it — diffusion_throwaway · 2026-07-27
- AI short film sets a 1985 dystopia to music and leans into cinema — ProfessorKey98 · 2026-07-27
- A new BOTPD episode made with Google Omni turns into an AI chase-scene parody — ScriptLurker · 2026-07-27
- A new LoRA recreates GTA: San Andreas’ classic RenderWare-era visuals — Humble-Pick7172 · 2026-07-27
- Enabling dynamic VRAM cuts LTX 2.3 video generation to 168s on an AMD R9700 — xdcfret1 · 2026-07-27