Speakrail: a fully-local open-source full-duplex voice assistant running on a single RTX 4090

danil_rootint · reddit · 2026-10-05

A developer open-sourced Speakrail, a fully-local, low-latency full-duplex voice agent that runs on one RTX 4090 and rivals GPT-Live on some benchmarks. It combines Voxtral Realtime with a turn-taking head, a microturn-finetuned Gemma 4 12B, and Breeze TTS 2, recreating the DuplexCascade approach with better models to add interruptions, interjections, and backchannels à la Thinking Machines' demo. Key lessons from iterating: early models wouldn't stop talking until prosody data from Voxtral was added; helper tokens derived from an MLP enabled turn-taking; and the smoking gun for missing interrupt behavior was that only 150 of 150k training samples were interruption examples. Full technical report to come.

Original post →

More from coding & agent

coding & agent channel →