Decagon: Detecting Relevant Speaker Changes by Combining Speaker Embeddings with Audio-Native Models

Scobleizer · x · 2026-08-28

Voice agents in customer support face a real-world mess: background voices, and users talking to other people at the same time. The agent must know when a different speaker actually joins the conversation, not react to every background voice.

Decagon's approach combines speaker embeddings with a post-trained audio-language model to determine when a speaker change is relevant. From Decagon's engineering blog.

Original post →

More from coding & agent

coding & agent channel →