Pruned CTC Cuts ASR Training Memory 5.1x, Enabling Native-LLM-Vocab Speech Recognition
X-LANCE · hf · 2026-09-29
X-LANCE introduces Pruned CTC, solving the memory explosion of CTC training with native LLM vocabularies.
Core observation & method
- Every valid CTC alignment uses only target tokens and blank; their union across a batch is a small vocabulary subset. Restricting alignment computation to this subset while keeping full-vocabulary normalization is provably exactly equivalent in loss and gradients
- Head-and-loss activation memory no longer scales linearly with vocabulary size; finite-beam alignment pruning is added
LLM-CTC
- Adapts pretrained LLMs for non-autoregressive ASR with causal attention and native vocabularies, extended to bounded-history streaming without chunk-level speech-text alignments
Results
- With Zipformer-M and 180K vocab: 5.1x memory reduction at only 17% step-time overhead, matching standard CTC accuracy
- On GigaSpeech across six Qwen3 sizes (0.6B–32B), LLM-CTC stays within 7% relative WER of LLM-CE with 7–10x faster recognition
- Fine-tuning Qwen3-ASR for streaming stays within 3% relative WER of matched offline models
More from Multimodal
- The more Suno you hear, the less you notice its flaws — context flooding in human-AI collaboration — voooooogel · 2026-09-29
- Testing Dreamina video generation with timestamped prompts and single-image reference — gen_ericai · 2026-09-29
- Testing Timestamped Prompts with Single Image Reference in dreaminacpp — gen_ericai · 2026-09-29
- Kling 4.0 Omni Reference demoed: multi-reference scene generation with instant language swaps — SarahAnnabels · 2026-09-29
- BananaStudio: Full-Resolution Nano Banana Image Studio Inside Gemini Canvas — Z3ROCOOL22 · 2026-09-29
- Suno Studio demo turns your voice into any instrument, starting with electric guitar — suno · 2026-09-29