Tailored ASR for Japanese speaking assessment cuts mora error rate from 12.3% to 7.1%
tkasasagi · x · 2026-09-27
Presented at Interspeech in Sydney, the paper "Building Tailored Speech Recognizers for Japanese Speaking Assessment" targets pronunciation-loyal ASR for Japanese-as-L2 education, outputting phonemic labels with accent markers.
Key points:
- Accurate accent-annotated phonemic transcriptions are scarce even for resource-rich Japanese.
- Two methods mitigate data sparsity: multitask training with auxiliary losses for orthographic text and pitch patterns, plus a finite-state transducer-based fusion of phonetic-alphabet and text-token estimators.
- On CSJ core evaluation sets, average mora-label error rate dropped from 12.3% to 7.1%, outperforming generic multilingual recognizers.
More from Multimodal
- Creator turns 1,000+ Midjourney images into animation using Claude Opus — ciguleva · 2026-09-27
- PrunaAI's distilled Qwen-Image-2.1 with few-step generation trends on Hugging Face — PrunaAI · 2026-09-27
- Dev's tested AI music workflow: lyrics, Suno, ear-curation, then Ableton — ctjlewis · 2026-09-27
- A full AI music video now costs ~$65 and 6M tokens — and it's no longer special — rickasaurus · 2026-09-27
- One Prompt, a 60-Second Singularity Video Essay: Runway CEO Demos Agentic Video Editing — c_valenzuelab · 2026-09-27
- Redditor argues Krea 2 is still the best full HD image model, ahead of its time — Due_Research9042 · 2026-09-27