Full-duplex voice models still fall short on humanlike zero-overlap turn-taking

alexisgallagher · x · 2026-09-13

The author notes that full-duplex audio models (like the one in the ChatGPT app) handle backchanneling more organically, but still don't achieve humanlike zero-overlap turn-taking, and none have shown convincing multi-speaker behavior.

Citing a recent benchmark paper: today's best turn predictors have poor recall (missing many genuine turn ends) and high false positives (mistaking pauses for turn ends). Worse, in real human conversation the turn gap is often negative — overlaps mean there is no gap at all — a fundamental challenge for current systems.

Related event: Study: voice AI turn-taking still far from human-level(3 posts)→

Original post →

More from Models

Models channel →