Open-Source MOSS Achieves Ultra-Low-Cost Audio Transcription and Separation

vanstriendaniel · x · 2026-07-10

Open MOSS has released MOSS-Transcribe-Diarize, a 0.9B parameter open-source model (Apache 2.0) capable of performing speech transcription, speaker diarization, and timestamping in a single pass. Developers used the model to process 174 hours of raw Apollo 11 mission audio, generating 45,000 timestamped speech segments. The entire process was completed on the Hugging Face platform using an A100 GPU in 3.8 hours, achieving a 47x real-time speed at a total cost of just about $9.46.

Related event: MOSS Releases Open-Source 0.9B Audio Transcription Model Supported by vLLM(5 posts)→

Original post →

More from Multimodal

Multimodal channel →