Surge AI: Post-Training on Office Work Boosts SWE-Bench Pro by 5.7pp, Zero Coding Data

echen · x · 2026-09-26

Surge AI post-trained Qwen3.5-122B-A10B on RL environments for long-horizon office work — documents, spreadsheets, web research, planning, tool use — with zero coding tasks, yet the model improved 5.7 percentage points on SWE-Bench Pro and gains generalized to unseen tool-use benchmarks.

The team attributes this to a general skill they call Goal-Directed Execution: forming precise goals, keeping them stable under pressure, maintaining an accurate picture of the environment, and verifying completion. Since every agent runs the same underlying loop, well-designed data can teach general capabilities rather than domain knowledge alone.

Original post →

More from Research

Research channel →