DAI-S2S-ST: 153K Human Preference Ratings Rank Voice Models on How They Feel to Talk To
rdesh26 · x · 2026-09-16
David AI released DAI-S2S-ST, a single-turn speech-to-speech human preference leaderboard built on 153K comparative human ratings across 7 models, 819 recorded prompts, and 12 evaluation dimensions from 283 raters.
The thesis: existing voice benchmarks measure task completion, but as people spend hours daily talking to models, naturalness, personality, and empathy — whether people want to keep talking — matter most. Initial results diverge from existing S2S leaderboards.
More from Models
- Speculation mounts OpenAI's mysterious 'new model' is a fresh pretrain, not an RL run — teortaxesTex · 2026-09-16
- GPT-6 Astra hits 68.7% on DrugDiscoveryBench, benchmark authors call it a step function — KexinHuang5 · 2026-09-16
- Zero scores 2.5% on Grade School Math vs base model's 62.2% — creators say it's not an assistant — maxsloef · 2026-09-16
- ChatGPT generates suggestive image, then refuses to swap its wolves for cats — comFX87 · 2026-09-16
- Blogger reviews Meta-linked Muse: 'violently poor design' but fast model and generous free tier — _AustinCalvert_ · 2026-09-16
- Virology researcher slams Claude's overzealous bio-safety filters while DeepSeek V4.1 just answers — Qwen30bEnjoyer · 2026-09-16