HiPLEX Factorizes Full-Duplex Speech Models Into Timing and Content Policies to Cut Interruptions
Qualcomm-AI-Research · hf · 2026-10-07
Qualcomm AI Research introduces HiPLEX, an RL framework for full-duplex speech language models that jointly optimizes turn-taking timing and semantic content, which prior methods optimized separately.
Design
- Factorizes a pretrained full-duplex text policy into a control policy (choosing pad/epad/con) and a conditional content policy (selecting tokens when 'con' is chosen), reusing the model's existing text head
- Routes timing advantages via event-causal masks from generated speech episodes, and an LLM-judge semantic advantage to the content factor
Results: Across three Moshi seeds on Full-Duplex-Bench v1, HiPLEX reduces takeovers during user pauses and backchannel opportunities and shortens post-interruption latency vs GRPO, while matching response quality; on Moshi and PersonaPlex it better matches pooled human turn-timing and backchannel-rate marginals.
More from Research
- WildMatch: weakly supervised matcher adaptation boosts wildlife re-identification accuracy — Turhan Can Kargin · 2026-10-07
- CTP: single-pass multimodal robot policy hits 97.25% on LIBERO, cuts inference latency to 75.8ms — Di Wu · 2026-10-07
- EAPN: execution-aligned noise fixes mode switching in asynchronous replanning, 96.7% bimanual success — Di Wu · 2026-10-07
- Two-step flow denoising cuts VLA inference from 61.6ms to 22ms for real-time robot control — Di Wu · 2026-10-07
- Magic-W0: a structured world-action foundation model topping RoboDojo-Sim at 27.10 — Xuhua Chen · 2026-10-07
- 'It Wanted To' Is the Researcher Talking: Pushing Back on AI Self-Awareness in Evals — gerardsans · 2026-10-07