Surge AI: Post-Training on Office Work Boosts SWE-Bench Pro by 5.7pp, Zero Coding Data
echen · x · 2026-09-26
Surge AI post-trained Qwen3.5-122B-A10B on RL environments for long-horizon office work — documents, spreadsheets, web research, planning, tool use — with zero coding tasks, yet the model improved 5.7 percentage points on SWE-Bench Pro and gains generalized to unseen tool-use benchmarks.
The team attributes this to a general skill they call Goal-Directed Execution: forming precise goals, keeping them stable under pressure, maintaining an accurate picture of the environment, and verifying completion. Since every agent runs the same underlying loop, well-designed data can teach general capabilities rather than domain knowledge alone.
More from Research
- Lukas Kaiser explains the recipe: distilled reasoning traces plus RL — lukaszkaiser · 2026-09-26
- François Fleuret explains LLMs: pretraining mimics humans, post-training steers toward valid answers — francoisfleuret · 2026-09-26
- wc3env: Warcraft 3 Frozen Throne turned into an open-source RL environment — daveholtz · 2026-09-26
- Cell-level gene expression links extend GWAS interpretation to the brain, new paper shows — anshulkundaje · 2026-09-26
- NeurIPS paper: sub-1% targeted perturbations can flip LMArena's top-ranked model — lintool · 2026-09-26
- Wild ICML/NeurIPS score swings fuel calls for peer review's demise amid LLM reviews — tokenbender · 2026-09-26