Open-Source Multi-Speaker Long Audio Transcription Model
dotey · x · 2026-07-19
OpenMOSS has released MOSS-Transcribe-Diarize-0.9B, focusing on multi-speaker long audio transcription and speaker diarization, aiming to answer "what was said, who said it, and when" in a single pass.
Key highlights include:
- 0.9B parameters, 128K context, capable of processing up to 90 minutes of audio at once without splitting or splicing;
- On a single 4090 GPU, achieves an RTF of 0.017, taking about 30 seconds to transcribe a 5–10 minute recording;
- Architecturally combines a Whisper-Medium encoder + Qwen3-0.6B style decoder, unifying ASR, speaker diarization, and time alignment into an end-to-end autoregressive task;
- Supports hotword enhancement and is open-sourced under Apache 2.0 for commercial use.
The post also includes links to the model and code.
Related event: OpenMOSS Releases MOSS-TD for Multi-Speaker Audio Transcription(2 posts)→
More from Models
- Google says Gemini 3.5 Pro is in partner testing as Gemini 4 pre-training starts — haider1 · 2026-07-22
- A benchmark chart puts a flash model around 5th place, but critics say it is far pricier — soumitrashukla9 · 2026-07-22
- How to Distinguish Genuine Token Efficiency from Shorter, Omissive Answers? — ruthstarkman · 2026-07-22
- Google reportedly ships Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber — gaganghotra_ · 2026-07-22
- China’s AI arms race is increasingly defined by chips, data centers, and open models — BenBajarin · 2026-07-22
- Sam Altman is headed to Washington to brief Congress on OpenAI’s GPT-6 line — inductionheads · 2026-07-22