PUMBA Paper Aligns Masked Diffusion LM Training With Inference Trajectories
kastnerkyle · x · 2026-10-07
PUMBA is a new paper addressing a core mismatch in masked diffusion language models: they are trained one way but sampled another. The authors study making MDM training more aware of the trajectories followed at inference time, which leads to more efficient text generation. Paper link in thread.
More from Research
- kalomaze on long-sequence training: conditional uncertainty and information asymmetry explain gradient absorption — kalomaze · 2026-10-07
- Ai2 publishes technical report on supercharging Olmo-core for scalable MoE training — StasBekman · 2026-10-07
- vf3 fuzzer unveiled at OAIC claims to outpace Jackalope and libprotobuf-mutator — dyn___ · 2026-10-07
- Scott Alexander's open letter to Steven Pinker: g-factor is real and AI scaling will keep climbing — Astral Codex Ten · 2026-10-07
- Paper shows diffusion transformer tokens encode lots of image info before it's interpretable — kwangmoo_yi · 2026-10-07
- Why synthetic cells die after five generations: they can't recycle their own broken parts — NikoMcCarty · 2026-10-07