PolyAI Releases Dialog-RSN-1: An End-to-End Speech-to-Speech Model

matthen2 · x · 2026-07-30

PolyAI introduces Dialog-RSN-1, an end-to-end audio LLM designed for speech-to-speech interactions. It integrates turn-taking, reasoning, response generation, and transcription within a single architecture.

By predicting whether it should speak as its very first token, the model achieves promptable turn-taking, highly accurate barge-in, and ultra-fast responses. It can proactively suggest moving to a quieter place by hearing background noise, and is capable of identifying emotions, accents, and talking speeds.

Related event: PolyAI Launches Dialog-RSN-1: End-to-End Native Audio Conversational Model(3 posts)→

Original post →

More from Models

Models channel →