Microsoft's ActiveSaddler Uses Automated Curriculum Learning to Boost Agent Harnesses by 7.5 Points

microsoft · hf · 2026-10-02

Microsoft introduces ActiveSaddler, framing training-scenario selection as the missing dimension of automated agent harness optimization. Existing methods update prompts, tool interfaces, and control logic from execution feedback while fixing the scenario curriculum; ActiveSaddler instead models the evolving curriculum as a non-stationary bandit with dynamically instantiated failure-pattern arms, balancing revisiting known weaknesses with exploring unseen scenarios. On GAIA2 and Terminal-Bench 2.0 it improves test Pass@1 by 4.4 and 7.5 percentage points over a fixed-order curriculum, with ablations confirming the gains come from dynamic target construction and utility estimation.

Original post →

More from coding & agent

coding & agent channel →