RL lifts Qwen 397B agent Pass@1 ~70% on APEX-Agents: SkyRL x Mercor full training postmortem

青稞AI · wechat · 2026-09-18

Edward Hu (first author of LoRA, ex-OpenAI o1 core member) joined Mercor as Head of AI Modeling and partnered with the UC Berkeley SkyRL team on RL for complex knowledge-work agents. Qwen3.5-397B-A17B improved Pass@1 on APEX-Agents from 16.11% to 27.29% (70% relative), while the smaller Qwen3.6-35B-A3B beat Claude Opus 4.5 after RL. Key engineering lessons: harness-only fixes added 5.95 points (22.74%→28.69%), TITO ensures inference/training token consistency to avoid off-policy drift, and fully async RL plus dynamic micro-batching maximizes throughput on long trajectories.

Original post →

More from coding & agent

coding & agent channel →