Open-Source MOSS Achieves Ultra-Low-Cost Audio Transcription and Separation
vanstriendaniel · x · 2026-07-10
Open MOSS has released MOSS-Transcribe-Diarize, a 0.9B parameter open-source model (Apache 2.0) capable of performing speech transcription, speaker diarization, and timestamping in a single pass. Developers used the model to process 174 hours of raw Apollo 11 mission audio, generating 45,000 timestamped speech segments. The entire process was completed on the Hugging Face platform using an A100 GPU in 3.8 hours, achieving a 47x real-time speed at a total cost of just about $9.46.
Related event: MOSS Releases Open-Source 0.9B Audio Transcription Model Supported by vLLM(5 posts)→
More from Multimodal
- Gemini Omni Flash turns a boat cabin into a cave in Flow by Google — chrisfirst · 2026-07-22
- A simple workflow to turn a photo into an image prompt using Gemini, Grok, or GPT Image — harshitagu72595 · 2026-07-22
- A Reddit user proposes a consistency LoRA to keep anime and game scenes visually stable — ThirdWorldBoy21 · 2026-07-22
- Hand-painted figurines run through Seedance look eerily alive — cocktailpeanut · 2026-07-22
- An AI agent-made bayou country music video is making the rounds on Reddit — LazyKaleidoscope4696 · 2026-07-22
- Testing Qwen 3 Image: Map Borders Shift Based on Prompts, Includes Chinese Labels — NirantK · 2026-07-22