Open-Source Audio Model Searches Apollo 11 Recordings
BanghuaZ · x · 2026-07-12
Open_MOSS has released **MOSS-Transcribe-Diarize**, an open-source **0.9B** model under the **Apache 2.0** license. This model combines **transcription**, **diarization**, and **timestamping** into a single pass, with its core advantage being suitability for **long audio inputs**. The provided real-world test case involves: - Processing **174 hours** of **Apollo 11** mission audio into searchable content - A total cost of just **$9.46** - Using original NASA recordings preserved by the Internet Archive - Processing a total of **103** audio segments - Completion via a single Hugging Face Job in **3.8 hours**, at a speed of roughly **47x realtime** - Generating **45,355** timestamped speaker segments The author notes that users can now search these mission recordings by "who said what and when."
Related event: Open_MOSS Releases Apollo Audio Transcription Model(2 posts)→
More from Multimodal
- Anatomy of Dynamic AI Images: Subject, Environment, and Camera — GPU_FieldNotes · 2026-07-21
- MiniCPM-V 4.6 now runs locally on iPhone with no cloud dependency — amos_gyamfi · 2026-07-21
- Creator says they no longer shoot with a camera, but with prompts — taherdhanera · 2026-07-21
- PixVerse demo turns into a full sci-fi dark comedy set on Mars — aliscodes · 2026-07-21
- Alibaba’s Qwen-Audio-3.0-TTS-Plus takes #1 on Artificial Analysis Speech Arena — airesearch12 · 2026-07-21
- Qwen Image 3 adds image editing and can generate multiple outputs in one pass — Linkpharm2 · 2026-07-21