Re-transcribing Seconds Around Each Cut Found Bad Seams That Whole-File ASR Missed
ustype · reddit · 2026-10-06
A developer building an agent that edits talking-head video shared a counterintuitive verification lesson: re-running ASR over the finished cut diffed against the script passed almost everything, because whole-file Whisper passes smooth over doubled words and phrases delivered twice from different takes — exactly the defects being hunted.
What works: (1) re-transcribe short windows aligned to each seam, not the whole file; (2) flag adjacent repeated words and clauses starting or ending mid-thought; (3) confirm a suspected double by re-transcribing just the 2–3 seconds straddling that seam, since overlapping windows can hallucinate repeats themselves; (4) never auto-fix a repeat — some languages (Urdu) double words on purpose, so the agent checks against the source. The broader design principle: the model never watches the video; audio is the clock, scripts handle mechanical work (silence detection, timing, ffmpeg) and the LLM only makes judgment calls. The project is open-sourced as an MIT-licensed local whisper.cpp skill at github.com/ranahaani/i-hate-editing.
More from coding & agent
- Yacine's 90-Minute Chat With a1zhang: Why Harnesses Boost LLM Generalization — yacinelearning · 2026-10-06
- Claude Opus 5.5 took 90,000 screenshots of its own game over two days to polish the visuals — prasenx · 2026-10-06
- classif: shell scripts branch on meaning via one token's logprobs on a local 12B — piotr1215 · 2026-10-06
- Agent 3D-renders its own game assets to build Star Control Hyper Melee unsupervised — draginol · 2026-10-06
- Markov: a pi-style, transparent LLM harness in a single Bash script with sub-agents — biller23 · 2026-10-06
- Meta paper: dual coding agents cross-review lift correct patches from 45.8% to 62.5% — rohanpaul_ai · 2026-10-06