MOSS Releases Open-Source 0.9B Audio Transcription Model Supported by vLLM

OpenMOSS-Team released MOSS-Transcribe-Diarize-0.9B, an open-source end-to-end audio understanding model with only 0.9 billion parameters. Capable of performing speech transcription, speaker diarization, timestamping, and acoustic event recognition in a single pass, the model gained day-one support from vLLM and quickly captured the open-source community's attention.

Key Features

Licensed under Apache 2.0, MOSS-Transcribe-Diarize-0.9B operates as an end-to-end audio-to-text pipeline. It effectively simplifies traditional complex speech processing workflows by supporting multi-speaker identification and long audio processing natively.

Practical Application and Cost Efficiency

Developer @vanstriendaniel tested the model's high efficiency and low-cost advantages. By utilizing this 0.9B model, he successfully processed 174 hours of original Apollo 11 mission audio from the Internet Archive, automatically generating text with timestamps and speaker labels for a total compute cost of less than $10.

2026-07-09 ~ 2026-07-10 · 5 related posts

1 near-duplicate retellings: vanstriendaniel