ASR paper boosts German disfluency F1 from 10% to 79% with verbatim control
nyralabs · hf · 2026-07-22
This paper treats transcription style in ASR—verbatim versus intended—as a latent variable that causes instability, evaluation confounding, and poor word-level timing.
Using coverage-aware decoder task tokens trained on parallel verbatim/intended transcript pairs, the authors report a large zero-shot gain on German disfluency F1, from 10% to 79%, even with English-only training. With full English fine-tuning, they say the method beats all baselines on verbatim accuracy, disfluency detection, and intended-mode quality across both languages.
The paper also introduces supervised cross-attention fine-tuning for better timestamps on disfluent speech and proposes verbatimize, a task for generating and enriching speech corpora with high-quality canonical verbatim transcripts.
More from Research
- Seven local open VLMs were tested as a second reader, and Qwen3-VL 4B won — MaziyarPanahi · 2026-07-22
- Bengaluru workshop will break down how world models work and how to build with them — tushaarmehtaa · 2026-07-22
- Baidu’s Unlimited OCR uses R-SWA to parse dozens of pages with a 3B model — vista8 · 2026-07-22
- Baidu’s Unlimited OCR uses R-SWA to parse long documents in one pass — vista8 · 2026-07-22
- Proposal maps Hermes Agent refactor to RIA and Logic Bus rules — Promptmethus · 2026-07-22
- YC-backed founder proposes Jcode bench v1 to make coding evals harder to game — Medium_Anxiety_8143 · 2026-07-22