Apple Proposes DACA-GRPO: Denoising-Aware Credit Assignment for Diffusion LLM RL
Apple ML Research · rss · 2026-09-16
Apple ML Research introduces DACA-GRPO (Denoising-Aware Credit Assignment for GRPO), targeting two fundamental weaknesses of RL for diffusion LLMs:
- Problem: existing methods treat all denoising steps as equally important, lacking temporal credit assignment across the denoising trajectory, and rely on biased, high-variance mean-field likelihood estimates
- Solution: DACA-GRPO is a lightweight, plug-and-play enhancement for any GRPO-style trainer, adding denoising-aware credit assignment
A method-level advance for RL fine-tuning of diffusion language models.
More from Research
- OpenAI Foundation funds "Public Data for Health" to create biology datasets for AI — DrMorganLevine · 2026-09-17
- New Philosophy of Science paper simulates AI-driven epistemic monocultures in research — MohammadAtari90 · 2026-09-17
- Harvard study suggests smartphones could predict suicide risk with striking accuracy — pshrink · 2026-09-17
- Study of 7 models across Claude Code, Codex, Pi: harness barely affects success but swings cost — DavideCrapis · 2026-09-17
- Neuroevolved value function solves Tetris at world-record speed in your browser — NathanWilbanks_ · 2026-09-17
- Polymarket atomicity gap: 1.8M reverted trades expose ghost-filled order attack — chaumian · 2026-09-17