Diffusion models can backprop end-to-end but costly; policy gradients only necessary for long-horizon language models
kalomaze · x · 2026-08-02
Author kalomaze notes that diffusion models can in principle backprop through the entire chain differentiably, but it's expensive. This cost forces the branching rollouts abstraction from ByteDance papers, which is only necessary for long-horizon language models. He also observes that almost no one uses policy gradients on diffusion models, except possibly ByteDance's branchgrpo-style works with in-house VLM reward models.
Related event: Policy Gradient in Diffusion Models Sparks Debate(3 posts)→
More from Research
- New "Discovery Episode" Framework Measures AI Scientists by Real Research Cycles — 量子位 · 2026-08-24
- AI Claims Breakthrough on Erdős Problem Transcendence — inductionheads · 2026-08-24
- Stanford's LLM-as-a-Verifier Boosts DeepSeek Score to 88% on Terminal-Bench — Saboo_Shubham_ · 2026-08-24
- Heterogeneous Quantum Architecture Cuts Physical Qubit Needs 138x for Fault Tolerance — MJBiercuk · 2026-08-24
- InfinityEdit: Infinite Video Editing via Lightweight Adapter — Yunze Tong · 2026-08-24
- Tencent Benchmarks Hybrid-Thinking MLLMs for Response Alignment — tencent · 2026-08-24