GraphForge: Evidence-Graph Workspace Synthesis Lifts GDPVal +65.7 With Only 2,169 Trajectories

ustc-community · hf · 2026-10-02

GraphForge synthesizes verifiable working-agent tasks grounded in real files via evidence graphs: tasks and rubrics are derived from file relations, each criterion anchored to files needed to verify it, with rollout-based executability checks and a revision agent. Fine-tuning Qwen3.6-27B on 2,169 trajectories brings GDPVal to 1445.7 (+65.7) under OpenHands, and Workspace-Bench-Lite and SpreadsheetBench II to +7.7 and +13.7 under Claude Code. Rejection fine-tuning using evidence-anchored rubrics adds further gains. Data and models are released.

Original post →

More from coding & agent

coding & agent channel →