Local Transcription Experience with Qwen3 ASR

dotey · x · 2026-07-17

The author reflects on the pain points encountered with speech-to-text during subtitle translation:

Recently, they tested Qwen3 ASR and found the results excellent. Paired with Qwen3-ForcedAligner, word-level timestamps can be aligned perfectly. Furthermore, the 0.6b model is sufficient and consumes few local resources.

For speaker recognition, they mentioned the open-source solution Pyannote + WeSpeaker, though accuracy remains limited when multiple people speak simultaneously; combining this with Agents and context could yield better results.

For higher quality demands, cloud solutions like the Doubao Audio File Recognition Model 2.0 on Volcengine are also viable. The author believes its quality and speed are quite good, though it requires extra payment.

Original post →

More from coding & agent

coding & agent channel →