RL Training Guide: Qwen 3.5 397B Pass@1 Boosted to 27.3%
mariofilhoml · x · 2026-09-02
Mercor Research published a detailed guide on post-training Qwen 3.5 397B using DPPO for long-horizon knowledge work. The model's Pass@1 rate on the APEX-Agents benchmark increased from 16.11% to 27.29%. The post covers the often overlooked infrastructure and de-risking steps, and releases the final model weights, full training scripts, and evaluation traces.
More from coding & agent
- Integrating Grok Bot into Superhuman via MCP: A 4-layer architecture — bfrench · 2026-09-02
- Qwen3.8-Max-0902 tops CodeArena WebDev with record score and pricing — Alibaba_Qwen · 2026-09-02
- Agentic engineering is fundamentally constraint engineering — burny_tech · 2026-09-02
- Smarter models overthink everything, slowing down coding workflows — sandyyevans · 2026-09-02
- InternReviewer Uses RL to Improve Peer Review Accuracy — Shanghai-AI-Laboratory · 2026-09-02
- Control-data flow separation keeps prompt optimization from breaking multi-agent pipelines — UWaterloo · 2026-09-02