RL lifts Qwen 397B agent Pass@1 ~70% on APEX-Agents: SkyRL x Mercor full training postmortem
青稞AI · wechat · 2026-09-18
Edward Hu (first author of LoRA, ex-OpenAI o1 core member) joined Mercor as Head of AI Modeling and partnered with the UC Berkeley SkyRL team on RL for complex knowledge-work agents. Qwen3.5-397B-A17B improved Pass@1 on APEX-Agents from 16.11% to 27.29% (70% relative), while the smaller Qwen3.6-35B-A3B beat Claude Opus 4.5 after RL. Key engineering lessons: harness-only fixes added 5.95 points (22.74%→28.69%), TITO ensures inference/training token consistency to avoid off-policy drift, and fully async RL plus dynamic micro-batching maximizes throughput on long trajectories.
More from coding & agent
- LangChain's Jev targets constrained output for building agent harnesses — multiply_matrix · 2026-09-18
- NewEyes AI brings a hybrid on-device/cloud visual agent to Meta glasses — rohanpaul_ai · 2026-09-18
- Cursor Projects wins praise: kanban tracking, multi-model agents and 3-way bakeoffs — kieranklaassen · 2026-09-18
- OpenAI's Astra Struggles With Long-Term Architecture: Duplication, Overengineering, Bad Assumptions Linger — Substantial_Swan_144 · 2026-09-18
- Hang Ten Systems raises $53M more in seed funding, totaling $85M — vsikka · 2026-09-18
- Does AI coding actually become multiplayer? Shared Slack agents vs. supercharged solo devs — Telos_in_the_Void · 2026-09-18