dQwen3.5: hybrid-attention diffusion LMs hit same loss with half the tokens
burny_tech · x · 2026-09-20
- Paper "dQwen3.5: Hybrid-Attention Diffusion Language Models" (Anton Xue, Sujay Sanghavi, et al.) adapts hybrid attention-RNN language models into diffusion models.
- Result: reaches a given training loss with roughly half the tokens of a full-attention control while supporting strong parallel (non-autoregressive) decoding.
- Practical takeaway: adapting existing open hybrid weights is far cheaper than training diffusion LMs from scratch, relevant for low-latency inference work.
More from Models
- $20/mo OpenAI users can't even pick the new "Sol" model — Sauers_ · 2026-09-20
- Ternary 2-bit Bonsai-2-27B GGUF lands on Hugging Face trending — dealignai · 2026-09-20
- Speculation: China still distilling frontier models implies weights weren't stolen — menhguin · 2026-09-20
- $1.74 to ask 1,738 model combos the same math question — only 6 got it right — arthurcolle · 2026-09-20
- User asks ChatGPT for UI icons, gets generated porn instead — w__flac · 2026-09-20
- Codex users hit surprise pauses with 90% of quota left, suspecting compute throttling — kfountou · 2026-09-20