Building a video-editing MCP, dev finds STT models like Parakeet v3 too smart — where's literal transcription?

Zeeplankton · reddit · 2026-09-23

A developer building a video-editing MCP hit a snag: STT models like Parakeet v3 try to be too smart, guessing sentence boundaries and stripping out ums — exactly what video editing needs to keep. They're looking for literal transcription models, ideally ones that also emit tags like (laughs) to give context on speaking cadence, and are asking for alternatives.

Original post →

More from coding & agent

coding & agent channel →