CalibForge: Automating Agent Task Generation via Adversarial Solver Calibration

AweAI-Team · hf · 2026-08-07

Training terminal agents effectively requires tasks that are both executable and appropriately challenging. Standard executable validation proves feasibility but fails to reveal how a task behaves relative to a specific solver.

This paper introduces CalibForge, an autonomous terminal-task synthesis system that revises candidate tasks through adversarial solver calibration. It employs two main strategies:

Both strategies operationalize a "solver-relative learnable zone." Using CalibForge, researchers constructed 5,431 calibrated tasks. Ablations show these strategies yield more effective supervision than authoring or single-solver feedback alone. Models trained on this dataset reached 47.57% on Terminal-Bench 2.0, with maximum improvements of 27.68 and 30.04 points on SWE-bench Pro and Doc2Repo, respectively.

Original post →

More from coding & agent

coding & agent channel →