A naive video transcription fixer pipeline: extract audio+frames, ASR, then correct with a frontier model

capetorch · x · 2026-09-29

capetorch proposes a step-by-step approach for correcting video transcriptions: extract audio and a screenshot per sentence, transcribe with a SoTA model, then fix errors using a frontier model with the captured images as context. He asks whether this naive approach can reach 99% accuracy.

Related event: Community Shares Video Transcription Error-Correction Recipe(2 posts)→

Original post →

More from coding & agent

coding & agent channel →