ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization
Sungho Park, Wonjoong Kim, Jue Zhang, Wook-Shin Han, Pengfei Gao, Chanyoung Park, Yongqiang Yao, Rao Fu, Elsie Nallipogu, Qingwei Lin, Victor Rühle
cs.AI, cs.CL, cs.LG, cs.MA, cs.SE
2026-10-01
ActiveSaddler turns scenario selection for harness optimization into a bandit over failure patterns, adding 4.4/7.5 Pass@1 points on GAIA2 and Terminal-Bench 2.0 at equal budget.
How reliably an LLM agent completes multi-step work depends heavily on its harness: the prompts, tool interfaces, and runtime control logic wrapped around the model. Automated harness optimization (AHO) methods such as GEPA, Meta-Harness, and AutoSaddler execute the current harness on training scenarios, diagnose weaknesses from traces, patch, and iterate.
All of them optimize how the harness is updated. Which scenarios generate the feedback is fixed before optimization starts: the full training set or a pre-scheduled mini-batch order. But the harness keeps changing, and so does the value of each scenario. An unresolved failure deserves another repair round; a repaired one has nothing left to teach. New failure modes sit undiscovered in unseen scenarios while old ones lose their remaining value. A fixed order handles neither side: many failures observed during fixed-order training remain unresolved in the final harness (Figure 3b). Under a limited rollout budget, that waste hurts.
ActiveSaddler casts the choice of training scenarios as a non-stationary multi-armed bandit. The arms are neither task categories nor individual scenarios but failure patterns: diagnosed, plausibly repairable harness weaknesses that recur across scenarios and traces. The curriculum layer wraps a fixed optimizer (AutoSaddler in the main experiments) and never touches its update mechanism.
Why failure patterns rather than categories or scenarios? The ablations answer directly. Category arms are too coarse: in one GAIA2 case, a still-failing scenario was folded into a broad time-category arm; after sibling scenarios passed, the arm's severity fell from 0.62 to 0.12 and that scenario was never sampled again. Scenario arms are too fine: a single weakness fragments across up to 75 arms on GAIA2, more than twice the maximum of 35 failure-pattern arms, so failures rarely get revisited under a limited budget.
Both benchmarks use gpt-5.5 for the task agent and the optimizer (medium and xhigh reasoning), a shared rollout budget (1,400 rollouts on GAIA2, 490 on TB2), and the mean of three test-time executions:
| Method | GAIA2 Pass@1 | TB2 Pass@1 |
| Manual harness | 53.6 ± 1.1 | 64.2 ± 2.9 (Terminus 2) |
| GEPA | 54.2 ± 2.2 | 65.8 ± 5.2 |
| Meta-Harness | 54.2 ± 1.2 | 66.7 ± 5.2 |
| AutoSaddler (fixed random order) | 55.4 ± 1.2 | 72.5 ± 0.0 |
| Difficulty-ordered fixed curricula | 55.9 / 55.7 | 70.8 / 73.3 |
| ActiveSaddler | 59.8 ± 1.0 | 80.0 ± 2.5 |
Same optimizer, same budget: +4.4 and +7.5 points. Against difficulty-ordered curricula the gap reaches +9.2 points (TB2), and 80.0 beats the hand-engineered Terminus-KIRA (69.2) by 10.8 points.
All three components matter: swapping in category arms costs 3.0/10.8 points (GAIA2/TB2), scenario arms 3.6/6.7, removing the Arm Prioritizer 4.5/7.5, removing the Exploration Controller 4.5/8.3. Non-LLM substitutes also fall short on GAIA2: EMA failure-persistence scoring reaches 56.7% and count-based UCB-AIR exploration 55.6%, versus 59.8%.
The mechanism checks out. Failure-pattern arms aim iterations at still-failing scenarios more often than category and scenario arms by 16.5/35.3 points on GAIA2 and 29.9/23.8 on TB2; of failures seen during training, 75.0% are fixed by the final harness on GAIA2 versus 50.0% for AutoSaddler. The DRAW rate falls from 44% in the first half of optimization to 20% in the second, and unresolved arms are 2.0x (GAIA2) and 6.9x (TB2) more numerous at PULL than at DRAW decisions, while a fixed schedule sits near 1x, insensitive to state.
It transfers: layered onto GEPA, GAIA2 rises from 54.2% to 57.2%; a second independent run gives 61.6%, still above every fixed-order baseline. The curriculum layer raises optimizer-side cost per patch from $7.87 to $11.44 on GAIA2, but end-to-end it is cheaper: reaching 58.5% dev accuracy costs $298 versus $1,360 for AutoSaddler; on TB2, 78.9% at $128 versus $220.
Teams optimizing agents have focused on how to update: reflective evolution in GEPA, diagnose-and-patch in AutoSaddler. This paper shows that which scenarios to train on is worth several points on its own, and it plugs into a different optimizer. What it saves is rollouts, the expensive part: spending them on unresolved weaknesses beats a uniform pass over the training set. The same idea should transfer to other feedback-driven improvement loops such as tool-definition or skill-library optimization.
To be clear, this is an incremental improvement: it replaces no optimizer, it makes existing ones earn more per rollout.
The paper has no dedicated limitations section; the points below mix its own disclosures with gaps the experiments leave open.