SteerDuplex: full-duplex speech model gains 44.5pp in steerability, new SteerBench released
ScaleAI · hf · 2026-09-22
ScaleAI researchers introduce SteerDuplex, targeting steerability — the underexplored ability of full-duplex spoken dialogue models to shift tone, persona, speaking rate, and voice style on instruction.
- Built on Moshi with fine-tuning on natural and synthetic dialogues covering instruction following, vocal delivery, reasoning, and duplex interaction
- Two-stage RL with hybrid rewards: verifiable interaction checks plus judge-based semantic feedback
- Released SteerBench: 390 spoken prompts and 1,067 human-authored binary audio/text rubrics across tone, persona, style/accent, and speed/length
- Results: supervised training lifts audio-steering pass rate by 44.5 percentage points over the strongest open baseline; Audio MultiChallenge improves 7 points; RL raises clean interruption response from 72.5% to 82.5% and cuts synthetic pause barge-in from 26.5% to 9%
- Reward probes reveal reward hacking via incomplete responses — timing gains must be evaluated alongside completeness
Model and benchmark are open for systematic research on spoken steerability.
More from Research
- Category-aware expert RL framework hits 59% on SWE-bench Multilingual — Logics-MLLM · 2026-09-22
- Stanford's three-tier controller gets GPT-6 Astra to 79.13% on RoboMME with just 3.63 calls per episode — DJiafei · 2026-09-22
- Fruit-fly-inspired spiking network Spi-Fly learns smells with few samples and little memory — Brighter-Side-News · 2026-09-22
- Decentralized multi-humanoid transport: one policy, no comms, pinch-lift-move (decMHT) — tweetsatpreet · 2026-09-22
- Terminal Bench Mini: 14-Instance Subset Replicates Full Local LLM Leaderboard Rankings — asankhs · 2026-09-22
- Model grafting turns Qwen3.5-4B into a causal encoder-decoder, 3.7x faster at 128K prompts — asankhs · 2026-09-22