ComfyUI-QwenASR v1.1.0: official Qwen3-ASR models, smart ITN, and long-form forced alignment
Narrow-Particular202 · reddit · 2026-09-12
ComfyUI-QwenASR v1.1.0 is out, closing the gap between raw speech recognition and usable text/subtitles in ComfyUI:
- Transformers 5 + official native models: legacy checkpoints deprecated; now uses Qwen3-ASR-1.7B-hf, Qwen3-ASR-0.6B-hf, and Qwen3-ForcedAligner-0.6B-hf with transformers >= 5.13.0 — pure upstream PyTorch, lower VRAM and faster on NVIDIA GPUs and Apple Silicon.
- Production-ready ITN: spoken numbers, percentages, and decimals auto-converted to numerals; spaced acronyms merged ("A S R" → ASR); whitelist protection keeps idioms intact.
- Hot-reloadable custom dictionary: rules live in itnrules. for custom acronyms and brand names across English, Chinese, Japanese, Korean, French — effective on the next run, no restart.
- Three nodes: ASR for lightweight transcription, Subtitle for sentence chunking with one-click .srt export, and Forced Align for long-form audio (podcasts/lectures) using iterative speaking-rate windowing, with automatic transcribe-and-align when transcript is left empty.
Install steps and sample workflows in the GitHub README.
More from Multimodal
- AI Mourns Humanity in Dark-Comedy Short 'We Leave the Lights On', Made with Suno v6 — AIandDesign · 2026-09-12
- Turning a 3D white model into a commercial: GPT-6 Astra + Seedance 2.5 + CapCut workflow — HeyAmit_ · 2026-09-12
- Paper-Tearing Comparison Video Shows a Year of Video-Gen Physics Gains — sabage27 · 2026-09-12
- Removing "AI slop" hallmarks from images with a single prompt to Astra — floguo · 2026-09-12
- Nari Labs open-sources Qwen3-TTS 1.7B: 50ms latency, 10x cheaper than ElevenLabs — alexcovo_eth · 2026-09-12
- Short film 'The Fly' made in 4 hours on one RTX 5090 with ComfyUI, Minimax, Krea 2 and Suno — aurelm · 2026-09-12