Amazon's MILO auto-evolves agent harnesses, hitting 86.1% on Terminal-Bench 2.1
amazon · hf · 2026-10-02
Amazon introduces MILO (Meta-evolutionary Island Orchestration), a framework that automatically discovers agent harnesses — the outer systems controlling model execution and environment interaction, which strongly affect long-horizon performance but demand heavy manual tuning.
How it works:
- Hierarchical lineage memory over island-based trees, using rejected mutations as negative evidence;
- Per-island mutator agents that rewrite complete harnesses using global search history and parent-specific feedback;
- An orchestrator adapting search via lineage grafting, speciation, mutator reassignment and curriculum revision.
Results: Across Terminal-Bench 2.1, PaperBench and DeepSWE, MILO-discovered harnesses beat eight SOTA harnesses and six search methods. With Opus 4.8 it improves resolution over the initial harness by +12.0%, +28.3% and +10.3%; Terminal-Bench 2.1 reaches 86.1±2.0%, topping the official leaderboard (83.8±2.3%) while using 26% fewer tokens. It also advances best-known bounds on EinsteinArena open math problems.
More from coding & agent
- Claude Code cloud sandbox now ships with Nix support — sloppenheimer · 2026-10-02
- NVIDIA paper: a better judge lifts terminal agent success from 50% to 68% without retraining — rohanpaul_ai · 2026-10-02
- Neuro-Symbolic Computer Use: agents that turn execution experience into self-healing policies, claimed 99% cheaper — xwang_lk · 2026-10-02
- Stripe now pays gas fees for agent stablecoin payments over MPP — jeff_weinstein · 2026-10-02
- The Flag Game: a toy setting to study agent swarm dynamics and cooperation — Hidenori8Tanaka · 2026-10-02
- Coinbase Link ships API to prove agents act on behalf of verified users — jeff_weinstein · 2026-10-02