Decagon: Detecting Relevant Speaker Changes by Combining Speaker Embeddings with Audio-Native Models
Scobleizer · x · 2026-08-28
Voice agents in customer support face a real-world mess: background voices, and users talking to other people at the same time. The agent must know when a different speaker actually joins the conversation, not react to every background voice.
Decagon's approach combines speaker embeddings with a post-trained audio-language model to determine when a speaker change is relevant. From Decagon's engineering blog.
More from coding & agent
- MATM Framework: Multi-Agent Transactive Memory Enables Population-Level Trajectory Sharing — 841io · 2026-08-28
- On Information Sharing and Subtasks in Multi-Agent Systems — 841io · 2026-08-28
- When do multi-agent systems beat a single agent? Look to distributed AI research — 841io · 2026-08-28
- Agent App: Reimagining Agent Interaction with WebUI and Bidirectional Notifications — DisastrousRelief9343 · 2026-08-28
- Open-source Discord AI assistant Zauq: Multi-model routing & Docker sandbox — rar_file-exe · 2026-08-28
- Auto-setup MulticaAI workspace using Claude Code or Hermes Agent — jiayuan_jy · 2026-08-28