WEFT: Whole-System Tool-Use Post-Training Lifts a 14B Agent Model by 12 Points on Claw-Eval
nex-agi · hf · 2026-10-05
nex-agi introduces WEFT (Whole-system Evolution For Tool-use Post-training), arguing that scaling executable environments in isolation doesn't guarantee gains because reliable learning signals depend on coherent interactions across the full agentic system: environment, task, agent harness, and evaluator.
Three pillars:
- Scalable interaction-system construction across environment breadth, task complexity, and interaction diversity;
- Execution-driven self-evolution that uses execution traces and state evidence to attribute failures and revise the responsible components, validated with fresh rollouts;
- Stable post-training at scale via prefix-preserving sampling, atomic-turn credit assignment, and MegaMCP for isolated, recoverable state across concurrent rollouts.
Results: WEFT-8B/14B beat matched-size environment-scaling baselines on BFCL V4, τ²-Bench, and Claw-Eval; WEFT-14B improves over Agent-World-14B by 6.41, 2.23, and 12.27 points; WEFT-35B-A3B extends gains to long-horizon benchmarks like Toolathlon-Verified and AutomationBench.
More from coding & agent
- marimo Studio: one notebook, separate audience views, agent-safe presentation layer — pandeyparul · 2026-10-06
- AI software factories: developers state intent, agents handle build and deploy — Pavan_Belagatti · 2026-10-06
- Dev builds polished card animations through multiple rounds with Astra — Dimillian · 2026-10-06
- pg-jev open-source Postgres extension runs semantic filters and scoring inside SQL — Arindam_1729 · 2026-10-06
- Open-source Discord AI assistant Zauq v4 ships bounded agents, MCP and sandboxed code verification — rar_file-exe · 2026-10-06
- Orbio launches new agent tools: sandboxes, 24/7 servers, databases and email, paid in CREDIT — econoar · 2026-10-06