MOSS-TD transcribes 90-minute multi-speaker audio and tracks who spoke when

udmrzn · x · 2026-07-21

The article presents MOSS-TD, a speaker-aware ASR system for transcribing long recordings such as 90-minute multi-speaker audio.

Instead of only converting speech to text, the system also identifies who is speaking and when, making it useful for long, messy conversations where speaker attribution matters as much as transcription accuracy.

Related event: OpenMOSS Releases MOSS-TD for Multi-Speaker Audio Transcription(2 posts)→

Original post →

More from Research

Research channel →