Microsoft's ActiveSaddler adapts agent harness training scenarios, boosting Pass@1 by up to 7.5 points
dair_ai · x · 2026-10-04
A new paper from Microsoft and colleagues introduces ActiveSaddler, a method for optimizing agent harnesses.
- Problem: Current harness optimizers only change how the harness is updated while keeping training scenarios fixed, so feedback from tasks becomes uninformative as the harness improves.
- Method: ActiveSaddler groups recurring failures into failure patterns, treating each pattern as an arm in a non-stationary bandit. It tracks how much the harness still learns from each pattern and splits the budget between revisiting known weaknesses and discovering new ones, adapting scenarios alongside the harness.
- Results: With the same optimizer, test Pass@1 improves by 4.4 points on GAIA2 and 7.5 points on Terminal-Bench 2.0 versus a fixed scenario order.
More from coding & agent
- When AI agents should escalate to humans: 5 trigger conditions — goyalshaliniuk · 2026-10-04
- 5 ways to stop AI agents from getting stuck in loops and losing the goal — goyalshaliniuk · 2026-10-04
- Open-source Hermes Gadget SDK turns an ESP32 into a voice terminal for your Hermes Agent — Teknium · 2026-10-04
- First AI film that truly stuns: an award-grade short coded end-to-end by Opus 5.5 — Alternative-Duty-532 · 2026-10-04
- How do you stop an LLM from inventing answers when the tool call never ran? — Aggressive-Narwhal-3 · 2026-10-04
- Dev Builds Native Metal Renderer for Minecraft With Claude's Help, Hits 60fps With Shaders on Old MacBook — EddyVGG · 2026-10-04