Un-fusing a realtime voice stack (STT → LLM → TTS) cut costs 14x — and the real win was text-level guardrails

Cloudsurfer_90 · reddit · 2026-09-22

After months in production, a developer ripped out their fused realtime voice model and rebuilt it as three stages: STT, then an LLM call, then TTS.

Original post →

More from coding & agent

coding & agent channel →