DAI-S2S-ST: 153K Human Preference Ratings Rank Voice Models on How They Feel to Talk To

rdesh26 · x · 2026-09-16

David AI released DAI-S2S-ST, a single-turn speech-to-speech human preference leaderboard built on 153K comparative human ratings across 7 models, 819 recorded prompts, and 12 evaluation dimensions from 283 raters.

The thesis: existing voice benchmarks measure task completion, but as people spend hours daily talking to models, naturalness, personality, and empathy — whether people want to keep talking — matter most. Initial results diverge from existing S2S leaderboards.

Original post →

More from Models

Models channel →