ASR paper boosts German disfluency F1 from 10% to 79% with verbatim control

nyralabs · hf · 2026-07-22

This paper treats transcription style in ASR—verbatim versus intended—as a latent variable that causes instability, evaluation confounding, and poor word-level timing.

Using coverage-aware decoder task tokens trained on parallel verbatim/intended transcript pairs, the authors report a large zero-shot gain on German disfluency F1, from 10% to 79%, even with English-only training. With full English fine-tuning, they say the method beats all baselines on verbatim accuracy, disfluency detection, and intended-mode quality across both languages.

The paper also introduces supervised cross-attention fine-tuning for better timestamps on disfluent speech and proposes verbatimize, a task for generating and enriching speech corpora with high-quality canonical verbatim transcripts.

Original post →

More from Research

Research channel →