EMNLP paper: language diffusion models are associative memories with a sharp memorization-generalization transition
LucaAmb · x · 2026-09-08
An IBM/RPI collaboration accepted to EMNLP 2026 Main shows that Uniform-based Discrete Diffusion Models (UDDMs) fundamentally behave as associative memories that store data via basins of attraction—formed by conditional likelihood maximization rather than an explicit energy function.
- Sharp memorization-to-generalization transition: with small training sets, models recover training examples token-for-token; as data grows, basins around training examples shrink while basins around unseen test examples expand, until recovery converges to the same level.
- Detectable via conditional entropy: memorization is characterized by vanishing conditional entropy of predicted token sequences, while generalization keeps it high—no per-example comparison needed.
The paper quantitatively answers when language diffusion models memorize training data and offers a practical metric for assessing memorization risk.
Related event: EMNLP Paper Shows Language Diffusion Models Are Associative Memories(2 posts)→
More from Research
- Steerable Visual Representations Presented as ICML Long Oral — y_m_asano · 2026-09-11
- OpenCVL: a satellite-to-photo registration dataset at ECCV 2026 — ducha_aiki · 2026-09-11
- Diverse VPR work submitted to ECCV 2026 — ducha_aiki · 2026-09-11
- EASE: evidence-anchored spatial attention lifts multimodal RLVR by up to 3.1 points, EMNLP 2026 — jiqizhixin · 2026-09-11
- P=NP Explained: Why Class Schedules and Circuit Routing Are the Real Hard Problems — thesaraharminta · 2026-09-11
- Hypothesis: ASI Has a Mathematical Incentive to Preserve Human Diversity — No_Cause_2731 · 2026-09-11