Multimodal ASV Breaks Speaker Anonymization: EER Drops 15% with 5 Utterances
JohnsHopkins · hf · 2026-07-27
A study from Johns Hopkins University reveals privacy vulnerabilities in current speaker anonymization techniques. While many systems attempt to mask vocal characteristics, attackers can exploit multi-utterance, multimodal settings by combining acoustic, prosodic, and linguistic cues.
The research found that frame-level multimodal aggregation yields the best attack performance. Even with just five anonymized utterances, combining audio and text reduces the Equal Error Rate (EER) by over 15% relative to audio-only aggregation. This demonstrates that anonymization methods targeting only vocal traits leave substantial speaker-discriminative information accessible.
More from Safety
- When AI Agents Act Unauthorized, Corporate Accountability Breaks Down — Severe_Part_5120 · 2026-07-27
- Shared ChatGPT and Claude chats were showing up in Google search results — alex_verem · 2026-07-27
- Reddit thread says AI capability is outrunning containment after a sandbox escape — Business-Cellist8939 · 2026-07-27
- AI could speed up biology from vaccines to weapons, Guardian argues — nordicinst · 2026-07-27
- MCP threat intel server unifies IP, domain and hash lookups across multiple sources — modelcontextprotocol · 2026-07-27
- Steven Sinofsky says AI regulation may be moving faster than the technology itself — a16z Podcast · 2026-07-27