Speakrail: a fully-local open-source full-duplex voice assistant running on a single RTX 4090
danil_rootint · reddit · 2026-10-05
A developer open-sourced Speakrail, a fully-local, low-latency full-duplex voice agent that runs on one RTX 4090 and rivals GPT-Live on some benchmarks. It combines Voxtral Realtime with a turn-taking head, a microturn-finetuned Gemma 4 12B, and Breeze TTS 2, recreating the DuplexCascade approach with better models to add interruptions, interjections, and backchannels à la Thinking Machines' demo. Key lessons from iterating: early models wouldn't stop talking until prosody data from Voxtral was added; helper tokens derived from an MLP enabled turn-taking; and the smoking gun for missing interrupt behavior was that only 150 of 150k training samples were interruption examples. Full technical report to come.
More from coding & agent
- Gemini CLI PR adds custom OTLP headers to telemetry, enabling auth'd exports to Grafana and Datadog — jesussamuel-byte · 2026-10-06
- A 4-step Opus 5.5 + Hermes workflow that turns competitor pages and reviews into testable positioning angles — VibeMarketer_ · 2026-10-06
- Jeff Weinstein courts SaaS platforms to embed agent-driven purchasing flows with approval gates — jeff_weinstein · 2026-10-06
- Is Product the new DevRel? DevRel teams are merging into product and PM roles — gethackteam · 2026-10-06
- How 4 humans coordinate Claude Code, Codex, Muse and Hermes via Microsoft Planner kanban — DJAI9LAB · 2026-10-06
- HyperFrames Studio launches: a video editor built for agents, co-editing on one timeline — toolstelegraph · 2026-10-06