WEFT: Whole-System Tool-Use Post-Training Lifts a 14B Agent Model by 12 Points on Claw-Eval

nex-agi · hf · 2026-10-05

nex-agi introduces WEFT (Whole-system Evolution For Tool-use Post-training), arguing that scaling executable environments in isolation doesn't guarantee gains because reliable learning signals depend on coherent interactions across the full agentic system: environment, task, agent harness, and evaluator.

Three pillars:

Results: WEFT-8B/14B beat matched-size environment-scaling baselines on BFCL V4, τ²-Bench, and Claw-Eval; WEFT-14B improves over Agent-World-14B by 6.41, 2.23, and 12.27 points; WEFT-35B-A3B extends gains to long-horizon benchmarks like Toolathlon-Verified and AutomationBench.

Original post →

More from coding & agent

coding & agent channel →