Follow-up on async RL training curves: rising loss metrics may measure staleness

yuxiangw_cs · x · 2026-09-18

Following his speculation that the team uses DAPO or an extension—with DeepSWE bench and AutomationBench pass rates climbing—the author follows up on the anomalous loss curves, guessing that the steadily rising loss metrics may measure staleness from asynchronous updates rather than genuine loss degradation.

Related event: Researcher Speculates DAPO-Based RL Training Behind Rising Benchmarks(2 posts)→

Original post →

More from Research

Research channel →