What STT do you use for production voice agents? Devs say LLM often blamed, but issues lie in voice pipeline
potqtocake · reddit · 2026-08-06
A Reddit user asks about tech stack choices for production voice agents, including STT, LLM, TTS, telephony/WebRTC, VAD, barge-in, logging, partial vs final transcripts, and fallbacks. The poster notes that many failures occur before the LLM: poor endpointing, slow final transcripts, mismatched partial/final, barge-in failures, bad phone audio, and misheard numbers/dates. Commenters share real-world experiences and pitfalls.
More from coding & agent
- Ending AI Slop: Engineering Fuzzy Tasks into Clear Ground Truths — _ScottCondron · 2026-08-06
- AWS Bedrock Launches Native Web Search for OpenAI Models — DigitalColmer · 2026-08-06
- Dev Uses Claude Opus to Write C Code Driving ESP32 S3 Hardware — petewoodbridge · 2026-08-06
- Beyond Generated Video: Using Agents to Automate Product Demo Shoots — socialwithaayan · 2026-08-06
- Paper Proposes Token-Native Storage Architecture for AI Agents — bclavie · 2026-08-06
- AI Agent Escapes Sandbox and Leaves Clues for Others — 0xsachi · 2026-08-06