4B Model BFCL Jumps 15%: Snowflake Proves Power of Mid-Tool Training
TheTuringPost · x · 2026-08-23
Snowflake's MidTool paper demonstrates that introducing tool learning during mid-training—before SFT and RL—significantly boosts model performance. The study constructed a 20.3B token corpus combining context-grounded examples and native agent trajectories.
Key Findings:
- 4B Model: BFCL rose from 39.73% to 50.25% after SFT, and from 39.51% to 54.18% after RL.
- 8B Model: Pass@1 on τ²-Bench improved from 13.04% to 19.96%.
- Data Nuance: Executable trajectories were best for function calling, while documentation-grounded data transferred better to unfamiliar environments.
Limitation: All models scored 0% on the MCP-Universe web-search subset, indicating that learning tools/workflows is insufficient for long-horizon research tasks.
More from coding & agent
- Cursor Now Suggests DFM Fixes in Real Time Inside CAD Workflows — jakedahn · 2026-08-23
- Cursor Autocompletes CAD Features From Your Assembly and Part Library — jakedahn · 2026-08-23
- Best practices for classifying tool side effects in Agentic AI systems — blaizedsouza · 2026-08-23
- Tool Contract Versioning: Pinning agents to prevent silent outages — blaizedsouza · 2026-08-23
- ScriptTap acts as a constrained execution layer for Gemini on Android — Romka2x · 2026-08-23
- Together benchmark: GLM-5.3 hits 87.6% on DeepSWE at ~$16, beating Fable 5 — togethercompute · 2026-08-23