Diffusion LMs Match Autoregressive Baselines on Math and Code After Pre-training
ricklamers · x · 2026-10-05
- A thread fragment on diffusion language model (dLM) results: compared with discrete counterparts (masked dLMs) and autoregressive baselines, the model is competitive after pre-training on math and code benchmarks including GSM8K, Minerva, HumanEval, and MBPP.
- After SFT, it comes close on hard reasoning, with some gains in instruction following and general knowledge.
- ricklamers shared the results approvingly; this is part 6 of a 10-post thread.
More from Models
- Security-One: open-weight 27B model outputs probabilities for agent security decisions — huggingface · 2026-10-05
- Red Hat AI ships NVFP4 quantized Qwen3.8-Flash-Next: MoE experts in FP4, vLLM-ready — huggingface · 2026-10-05
- GPT-6.1 Sol tested on Terminal-Bench: xhigh is the sweet spot, medium degrades badly — aitrendz_xyz · 2026-10-05
- Opus 5.5 vs Sol: browser plush octopus with combable fur, Opus wins on fluff and price — aitrendz_xyz · 2026-10-05
- Claude Conversation Monitoring Sparks Backlash and Local AI Push — zacharynado · 2026-10-05
- OpenAI rolls out textGrain text watermarking for EU AI Act, going open source — btibor91 · 2026-10-05