Google's ToolGrad generates tool-use datasets answer-first, hitting near 100% pass rate
DuRuofei · x · 2026-09-11
Google, University of Tokyo, and RIKEN AIP researchers introduced ToolGrad, an efficient framework for generating LLM tool-use datasets, released at ACL 2026 Findings with open data and code.
- Key idea: inverts the usual pipeline — it first generates ground-truth tool-use chains using textual "gradients", then constructs prompts from them (answer-first instead of question-first).
- Results: near 100% pass rate for generated data, lower generation cost, and improved LLM tool-use performance after training.
- Data format: each sample includes a query, a tool-call graph, and step-by-step tool invocations, covering multi-tool scenarios like news search, company lookups, domain monitoring, and TikTok search.
Led by Zhongyi Zhou; arXiv paper, code, and dataset are all available.
More from coding & agent
- Bezalel launches one-URL agent platform bundling memory, email, money, texting, desktop, sandboxes and connectors — Rasmic · 2026-09-11
- Give a Spreadsheet Agent a Response Budget, Not Just a Cell-Range Argument — Medical-Cow289 · 2026-09-11
- Tsinghua's DiffuTester accelerates dLLM unit test generation up to 3x via AST pattern mining — 机器之心 · 2026-09-11
- Ten lessons from three years building agents for real production work — garrytan · 2026-09-11
- Devin's New Model Verdict: Not a Benchmaxxer, a 'Killer Execution Model' at $20/Month — brandon_galang · 2026-09-11
- Shopify CEO Tobi Lütke hails single-dev open-source agent harness Pi — aakashgupta · 2026-09-11