Duplex Cue: New Benchmark Tests Whether Voice Agents Adapt While Speaking
rdesh26 · x · 2026-09-15
The team behind Besimple AI released Duplex Cue, an audio benchmark for evaluating in-turn adaptation behavior in full-duplex voice agents, available on Hugging Face with an accompanying paper.
- Problem: Full-duplex evaluation typically reduces overlap handling to a binary keep-speaking-or-stop score, missing the third response humans use routinely: continuing to talk while incorporating what the listener just contributed.
- Design: The benchmark labels listener intent (backchannel, collaboration, or interruption) independently from speaker behavior (continued, adapted, yielded), producing a 3×3 matrix that makes in-turn uptake visible instead of treating every overlap as an interruption.
- What it measures: Whether an agent ignores a mid-turn cue, acknowledges or revises content on the fly while holding the floor, or properly hands over the turn.
More from Models
- Yandex open-sources its search AI answer model, squeezing 40% more answers from same compute — teortaxesTex · 2026-09-15
- Opus 5 reportedly routing to Opus 5.2 for some users — ResultBackground2450 · 2026-09-15
- Brockman claims HF-hack model lacked alignment training; report contradicts — sebkrier · 2026-09-15
- Resemble AI ships DETECT-World, a physics-based deepfake detector with 99.5% audio accuracy — AiBreakfast · 2026-09-15
- Grok 4.7 Misses Target Again; 2.5T-Parameter Grok 4.8 Finishes Training — eyishazyer · 2026-09-15
- Writers Report Gemini Flash 3.8 Suffers Severe Long-Context Rot Despite 1M Token Window — Quenty1 · 2026-09-15