Looking for a model to label each subtitle line with the speaker's identity
dtdisapointingresult · reddit · 2026-09-24
For a fan subtitling project, the author needs to convert unlabeled English TV subtitles into per-line speaker-attributed format. With 15 main characters, they're willing to prep one reference voice sample per actor (Alice.wav, Bob.wav) and want automatic speaker identification, marking unrecognized extras as Unknown #1/2/3 for manual cleanup (1% of lines). They ask whether an existing model can do reference-audio-based speaker labeling.
Related event: Subtitle maker seeks speaker diarization solution for 15-character show(2 posts)→
More from Multimodal
- Runway CEO declares video the final interface alongside real-time video UI demo — _AustinCalvert_ · 2026-09-24
- Gemini 3.8 Flash TTS sweeps all seven Voice Arena language boards to #1 — kastnerkyle · 2026-09-24
- jxnl's robust music transcription pipeline: Astra + spectrogram error correction — jxnlco · 2026-09-24
- AI-made music video on datacenter water use: GPT 6 + Suno + Slop Cannon pipeline — zealcaiden · 2026-09-24
- AI Demo Generates Explorable Mars Life, Takes Real-Time Scene Direction — karinanguyen · 2026-09-24
- ACTx486 demo lets you talk to any video and get real-time responses — karinanguyen · 2026-09-24