Full-sequence masking during SFT unlocks prompt infilling for diffusion LLMs, COLM 2026 paper finds

kastnerkyle · x · 2026-09-24

A Patronus AI paper accepted at a COLM 2026 workshop shows masked diffusion LLMs like LLaDA and Dream can't infill prompts only because SFT masks responses exclusively. Switching to full-sequence masking—masking prompts and responses jointly—lets models infer task-adapted prompts from few-shot examples that match or beat manual templates and transfer across models. The authors argue training practices, not architecture, are the bottleneck, and call on the community to release full-sequence-masked SFT checkpoints.

Original post →

More from Research

Research channel →