New SI-cpWER benchmark tests persistent speaker identity across meetings; ThyVoice leads at 47.13
Shantanu Vispute · hf · 2026-10-02
A new HF paper introduces Speaker Identified cpWER (SI-cpWER), scoring a corpus under one global speaker-ID assignment to measure whether the same person keeps one identity across meetings. Evaluated on the full 129-meeting CHiME-8 NOTSOFAR set plus CHiME-6 against five commercial cascades and two academic baselines, the end-to-end ThyVoice system (overlap repair plus gated voiceprint updates) beats every commercial system in all conditions, 47.13 vs 54.75 mean SI-cpWER for the runner-up. Persistent attribution must be evaluated directly in memory-reuse systems.
More from Research
- NVIDIA's Mid-Harness: a strong verifier boosts terminal agent Pass@1 from 50% to 68% on TerminalBench-Lite — rohanpaul_ai · 2026-10-02
- NVIDIA paper: a better judge lifts terminal agent success from 50% to 68% without retraining — rohanpaul_ai · 2026-10-02
- JevBench to add evals for LLM routing, RAG retrieval, and moderation use cases — airesearch12 · 2026-10-02
- Neuro-Symbolic Computer Use: agents that turn execution experience into self-healing policies, claimed 99% cheaper — xwang_lk · 2026-10-02
- The Flag Game: a toy setting to study agent swarm dynamics and cooperation — Hidenori8Tanaka · 2026-10-02
- How Matei Zaharia went from Berkeley PhD and Spark to a $190B Databricks — jfiance · 2026-10-02