RL Training Guide: Qwen 3.5 397B Pass@1 Boosted to 27.3%

mariofilhoml · x · 2026-09-02

Mercor Research published a detailed guide on post-training Qwen 3.5 397B using DPPO for long-horizon knowledge work. The model's Pass@1 rate on the APEX-Agents benchmark increased from 16.11% to 27.29%. The post covers the often overlooked infrastructure and de-risking steps, and releases the final model weights, full training scripts, and evaluation traces.

Original post →

More from coding & agent

coding & agent channel →