ToolGrad flips tool-use data synthesis: answer-first pipeline hits ~100% pass rate

tokenbender · x · 2026-09-16

ToolGrad (ACL 2026 Findings) inverts the standard tool-use dataset synthesis paradigm: instead of generating user queries first and annotating tool-use chains via DFS, it builds valid tool-use chains iteratively guided by textual "gradients," then synthesizes matching queries. The resulting ToolGrad-500 dataset has more complex tool use, lower cost, and near-100% pass rate, and models trained on it beat baselines trained on expensive datasets and proprietary LLMs. Code, data, and models are open-sourced. The poster highlights a reusable recipe: collect valid outcomes, derive likely goals, mutate for diversity, verify, train.

Related event: 'Answer-first' synthetic data method hits nearly 100% pass rate(2 posts)→

Original post →

More from Research

Research channel →