Building a video-editing MCP, dev finds STT models like Parakeet v3 too smart — where's literal transcription?
Zeeplankton · reddit · 2026-09-23
A developer building a video-editing MCP hit a snag: STT models like Parakeet v3 try to be too smart, guessing sentence boundaries and stripping out ums — exactly what video editing needs to keep. They're looking for literal transcription models, ideally ones that also emit tags like (laughs) to give context on speaking cadence, and are asking for alternatives.
More from coding & agent
- Sebastian Raschka: the real appeal of open-source agent harnesses is inspectability, not price — rasbt · 2026-09-23
- Cross-posting tool Ferryman nears $10,000 MRR with $30-$100/mo tiers and posting straight from Claude and Cursor — KevinNaughtonJr · 2026-09-23
- Swarms releases 5 ecosystem guides comparing its agent framework with CrewAI, LangGraph and Autogen — KyeGomezB · 2026-09-23
- 15-year ads veteran builds the ad-platform MCP he couldn't find: 14 sources, paused-by-default writes — DapperManagement1306 · 2026-09-23
- Critical Next.js RCE: CVE-2026-94545 hits next/og ImageResponse via Satori SVG escaping flaw — evilsocket · 2026-09-23
- Cognition floods Devin with GPT-6 models and cuts task costs 61%, gives away 50 Max plans — EricBuess · 2026-09-23