Follow-up on async RL training curves: rising loss metrics may measure staleness
yuxiangw_cs · x · 2026-09-18
Following his speculation that the team uses DAPO or an extension—with DeepSWE bench and AutomationBench pass rates climbing—the author follows up on the anomalous loss curves, guessing that the steadily rising loss metrics may measure staleness from asynchronous updates rather than genuine loss degradation.
Related event: Researcher Speculates DAPO-Based RL Training Behind Rising Benchmarks(2 posts)→
More from Research
- CMU and Oxford Paper Shows Looped Flows Hit 58.8% on ARC-AGI-1 Without Long Chains of Thought — rohanpaul_ai · 2026-09-18
- CMU and Oxford paper: refining hidden state beats longer chain-of-thought for reasoning — rohanpaul_ai · 2026-09-18
- An agent foundations reading list: tiling agents, FDT, logical induction and naturalized induction — jessi_cata · 2026-09-18
- RL agents invent their own diagnostic renderings to ground code understanding, sparking RL scaling optimism — teortaxesTex · 2026-09-18
- Grounded SI: egocentric data at scale hinges on hand tracking in wild, long-tail scenarios — micoolcho · 2026-09-18
- GenBio AI co-founders publish 'A world model of the virtual cell' in Cell — HongyiWang10 · 2026-09-18