Researcher Speculates DAPO-Based RL Training Behind Rising Benchmarks

A researcher infers from public training curves that a team is likely using DAPO or its extensions for RL, with pass rates climbing on DeepSWE and AutomationBench, and speculates that rising loss may reflect update lag in async training.

2026-09-18 ~ 2026-09-18 · 2 related posts