Duplex-MPE benchmarks multi-party full-duplex speech: MiniCPM-o 4.5 leads on three of four scores
PKU · hf · 2026-09-29
Duplex-MPE tests when a full-duplex speech assistant should answer, stay silent, or stop speaking in multi-party conversations. The benchmark has 2,000 scenarios with 3-4 human speakers plus an assistant, paired across explicit/implicit addressing, with continuous audio and no transcripts. Among five open-weight systems evaluated, MiniCPM-o 4.5 leads on three scored capabilities; frequent speech from others coexists with wrong answers or failure to stay silent. A transcript-based Gemini 3.1 Pro reference responds 64.3 points more to explicit than implicit requests, while speech systems show no significant difference — current models can't tell if they're being addressed.
More from Models
- Leaked OpenAI 'dot' details show raising phone to ear triggers ChatGPT Voice — koltregaskes · 2026-09-29
- Google to replace Gemini Gems with Skills starting November 17 — mark_k · 2026-09-29
- Carla v0.1.0: a local llama.cpp loom TUI for growing AI characters — max_paperclips · 2026-09-29
- Leaked OpenAI DevDay reveal called 'just a Grok bot / Meta Muse rip-off' — gaganghotra_ · 2026-09-29
- Running Jev at high frame rate with full-state snap inferences makes it a true System 1 — mathemagic1an · 2026-09-29
- OpenAI model naming rumor: Dots, Orbit and 'o' said to be in the mix — mark_k · 2026-09-29