GraphForge: Evidence-Graph Workspace Synthesis Lifts GDPVal +65.7 With Only 2,169 Trajectories
ustc-community · hf · 2026-10-02
GraphForge synthesizes verifiable working-agent tasks grounded in real files via evidence graphs: tasks and rubrics are derived from file relations, each criterion anchored to files needed to verify it, with rollout-based executability checks and a revision agent. Fine-tuning Qwen3.6-27B on 2,169 trajectories brings GDPVal to 1445.7 (+65.7) under OpenHands, and Workspace-Bench-Lite and SpreadsheetBench II to +7.7 and +13.7 under Claude Code. Rejection fine-tuning using evidence-anchored rubrics adds further gains. Data and models are released.
More from coding & agent
- Dev recreates a DOOM-like game with GPT, nailing the original's gory feel — DeryaTR_ · 2026-10-02
- Giving an AI agent real control of a homelab — responsibly — Academic_Wolverine · 2026-10-02
- Process-mining agents found 20 steps and 7 loops in a workflow documented as 7 steps — vasuman · 2026-10-02
- Geoffrey Huntley: Forget reading code—your verification properties are all that matters — kieranklaassen · 2026-10-02
- Claude Skills explained: why prewritten PDF scripts beat pasting prompts every time — lxfater · 2026-10-02
- Eye.Art Polyphemus: a chat-first MCP for image generation and reference-based edits — axiomofaxiom · 2026-10-02