Agent Training Data: More Tools, Harder to Generate
Shahules786 · x · 2026-07-14
This thread discusses how to generate post-training data that genuinely enhances agent capabilities. The core insight is that in enterprise-level, multi-turn long tasks, the number of tools expands exponentially. A single domain might have 100+ tools, and three domains easily exceed 300. Simply stuffing the tool list into the context degrades task performance.
The thread highlights three papers:
- PlanBench-XL: Constructs a tool environment closer to real-world scenarios, simulating agents working without seeing the full toolset;
- TMax: Focuses on synthesis and training for harder tasks;
- Autodata: Uses multi-model collaboration to determine if a task is feasible and trainable.
The author notes these papers are already influencing their internal practices: synthesizing tasks from tool graphs and combining multiple models of varying capabilities to assess trainability, with the ultimate goal of integrating these methods into a customizable, automated data pipeline.
Related event: PlanBench-XL Tackles Agent Tool Retrieval at Scale(3 posts)→
More from coding & agent
- MathModelAgent gains traction: auto-solves math modeling and writes a submission-ready paper — jihe520 · 2026-09-11
- alphaXiv open-sources OpenResearch to run parallel research agents with any model — alphaXiv · 2026-09-11
- DeskcommCRM: open-source AI sales CRM with native agents and WhatsApp hits 1k stars — melgarafael · 2026-09-11
- hyperresearch: agent-driven knowledge base that turns web research into a searchable wiki — jordan-gibbs · 2026-09-11
- Forter's 13 lessons from its agent sprint: skip custom RAG, lean on mature enterprise search — bibryam · 2026-09-11
- Two real 'company brains' opened up live: Gorgias' in-house Cortex vs Slite — femke_plantinga · 2026-09-11