Researcher speculates DAPO-based RL training as DeepSWE and AutomationBench pass rates keep climbing

yuxiangw_cs · x · 2026-09-18

Based on public training curves, the author speculates the team is likely using DAPO or an extension of it for RL training: pass rates on DeepSWE bench and AutomationBench keep climbing, and several loss metrics rise alongside them. He finds the rising losses counterintuitive and wonders whether they are cumulative or the rollout length is growing over training.

Related event: Researcher Speculates DAPO-Based RL Training Behind Rising Benchmarks(2 posts)→

Original post →

More from Research

Research channel →