MOSS Releases Open-Source 0.9B Audio Transcription Model Supported by vLLM
OpenMOSS-Team released MOSS-Transcribe-Diarize-0.9B, an open-source end-to-end audio understanding model with only 0.9 billion parameters. Capable of performing speech transcription, speaker diarization, timestamping, and acoustic event recognition in a single pass, the model gained day-one support from vLLM and quickly captured the open-source community's attention.
Key Features
Licensed under Apache 2.0, MOSS-Transcribe-Diarize-0.9B operates as an end-to-end audio-to-text pipeline. It effectively simplifies traditional complex speech processing workflows by supporting multi-speaker identification and long audio processing natively.
Practical Application and Cost Efficiency
Developer @vanstriendaniel tested the model's high efficiency and low-cost advantages. By utilizing this 0.9B model, he successfully processed 174 hours of original Apollo 11 mission audio from the Internet Archive, automatically generating text with timestamps and speaker labels for a total compute cost of less than $10.
2026-07-09 ~ 2026-07-10 · 5 related posts
- [source] MOSS Audio Transcription Model Integrates with vLLM — vllm_project · 2026-07-09
- MOSS Speech Transcription Model Goes Viral — OpenMOSS-Team · 2026-07-09
- [source] MOSS Releases Multi-Speaker Transcription Model — pmttyji · 2026-07-09
- Open-Source MOSS Achieves Ultra-Low-Cost Audio Transcription and Separation — vanstriendaniel · 2026-07-10
1 near-duplicate retellings: vanstriendaniel