A naive video transcription fixer pipeline: extract audio+frames, ASR, then correct with a frontier model
capetorch · x · 2026-09-29
capetorch proposes a step-by-step approach for correcting video transcriptions: extract audio and a screenshot per sentence, transcribe with a SoTA model, then fix errors using a frontier model with the captured images as context. He asks whether this naive approach can reach 99% accuracy.
Related event: Community Shares Video Transcription Error-Correction Recipe(2 posts)→
More from coding & agent
- Millions of person-hours wasted building AI harnesses, erased by new model releases — sebpaquet · 2026-09-29
- Unverified Claim: Anthropic Engineers Share All Claude Sessions, Teammates Can Steal Each Other's Tasks — YouJiacheng · 2026-09-29
- Open-source self-driving sim repo auto-galleries 38 demos, adds Claude Code PR review skill — 4310sy · 2026-09-29
- Celesto: open-source persistent microVM computers for AI agents, boots in 500ms — aniketmaurya · 2026-09-29
- A practical guide to adopting AI in your organization: management skills over prompt tricks — chribonn · 2026-09-29
- RSI Arena: 8 AI agents get 1,000 GPU-hours each to train a better model live — my_cat_can_code · 2026-09-29