CalibForge: Automating Agent Task Generation via Adversarial Solver Calibration
AweAI-Team · hf · 2026-08-07
Training terminal agents effectively requires tasks that are both executable and appropriately challenging. Standard executable validation proves feasibility but fails to reveal how a task behaves relative to a specific solver.
This paper introduces CalibForge, an autonomous terminal-task synthesis system that revises candidate tasks through adversarial solver calibration. It employs two main strategies:
- Multi-solver calibration: Targets disagreements within a heterogeneous pool of solvers.
- Contrastive solver calibration: Targets a specific strong-pass/weak-fail dynamic.
Both strategies operationalize a "solver-relative learnable zone." Using CalibForge, researchers constructed 5,431 calibrated tasks. Ablations show these strategies yield more effective supervision than authoring or single-solver feedback alone. Models trained on this dataset reached 47.57% on Terminal-Bench 2.0, with maximum improvements of 27.68 and 30.04 points on SWE-bench Pro and Doc2Repo, respectively.
More from coding & agent
- Claude Skills Masterclass: From Single Prompt to Full Repo and Beyond — aakashgupta · 2026-08-07
- Cloudflare Launches Kitesurf: A Headless Browser Built for AI Agents — dinasaur_404 · 2026-08-07
- Stripe hires engineers to build tools for AI agents to buy software autonomously — jeff_weinstein · 2026-08-07
- Seamless MiniMax H3 Video Chaining: Motion and Audio Continue Across Clips — Sad_Berry_4621 · 2026-08-07
- Extracting Meta Agent System Prompts Boosts Cline Bug-Fixing Efficiency — alejandroll10 · 2026-08-07
- Aeon: A Fully Autonomous Agent Framework Without Approval Loops — tom_doerr · 2026-08-07