Researcher speculates DAPO-based RL training as DeepSWE and AutomationBench pass rates keep climbing
yuxiangw_cs · x · 2026-09-18
Based on public training curves, the author speculates the team is likely using DAPO or an extension of it for RL training: pass rates on DeepSWE bench and AutomationBench keep climbing, and several loss metrics rise alongside them. He finds the rising losses counterintuitive and wonders whether they are cumulative or the rollout length is growing over training.
Related event: Researcher Speculates DAPO-Based RL Training Behind Rising Benchmarks(2 posts)→
More from Research
- Info geometry note: categorical distributions form both a mixture and exponential family — FrnkNlsn · 2026-09-18
- After months of work, team reconstructs 3D human-object motion from plain video — andrew_n_carr · 2026-09-18
- NGX1 uses AI and multivalent physics to deliver mRNA to any cell in the body — ycombinator · 2026-09-18
- Hydro merges Verus-checked commutativity proofs for distributed systems with zero hand-written specs — ShadajL · 2026-09-18
- Notes on all 13 lectures of Nathan Lambert's RLHF course: one scalar reward is the root of most complaints — le_james94 · 2026-09-18
- Fine-tuned personas all refuse unsafe requests, and eval scores belong to the harness, not the weights — le_james94 · 2026-09-18