LLaDA 2.2 uses L-EBPO to cut error propagation in long agent runs
omarsar0 · x · 2026-07-26
The post explains how LLaDA 2.2 tries to avoid long-horizon collapse during RL.
- Its RL stage builds on the same block-level editing mechanism used earlier in training.
- L-EBPO extends block-level policy optimization with editing operations, teaching the model when to cut a defective span and when to fill a gap from environment feedback.
- The claim is that this interrupts trajectory-level error propagation, which is a structural fix for collapse in long agent runs.
More from Models
- Arav Srinivas calls GLM underrated and says 700B feels close to Opus — AravSrinivas · 2026-07-28
- Kimi K3 reproduces RLVR findings without overclaiming, author says — infoxiao · 2026-07-28
- Inference.net pitches a gateway flow that mirrors prod traffic to Kimi K3 before switching — MatthewBerman · 2026-07-28
- Claude was unsubscribed as ChatGPT/Codex 5.6, Sol and Kimi 3 all struggled — sull · 2026-07-28
- GLM 5.5 is said to arrive in August with stronger long-horizon agent loops — bindureddy · 2026-07-28
- Moonshot’s Kimi K3 lands in Japan with 2.8T open weights and $3/$13 pricing — DavidBennett__ · 2026-07-28