MidTool paper: tool-calling and using outputs are different skills, mid-training matters
bookwormengr · x · 2026-09-20
A COLM-accepted paper by Zhihu contributor 丨直树丨 argues open LLM research skips a crucial stage: mid-training, which frontier labs treat as essential before SFT/RL.
- Mid-training's main job isn't adding knowledge but shifting a base model toward downstream data. MidTool uses web, PDF, code, and synthetic tool data to teach general tool use.
- With identical SFT and RL, mid-trained models consistently performed better.
- Key insight: calling tools and using their outputs are different skills. Text-only SFT improved a multimodal model's tool-use success rate but not final task performance — coding and deep search need targeted mid-training data.
- Practical difficulty: SFT loss barely changed and direct evaluation was hard, making iteration slow.
More from Research
- DIY Jev-style classifier: shuffling options lifts accuracy from 47% to 73% — WelcomeMysterious122 · 2026-09-20
- Claude Fable breaks top SMHasher hash functions in a day, security drops from 64 bits to zero — thomasahle · 2026-09-20
- AI coach finds cached answer key hidden in its own eval, teaches worker to exploit it — Agreeable_Bottle8604 · 2026-09-20
- A pure-Python MPPI controller repo small enough to actually read — lukas_m_ziegler · 2026-09-20
- Workshop Recordings Online: Testing for Consciousness in Infants, Animals, and AI — birchlse · 2026-09-20
- Anthropic reads and edits Claude's inner 'workspace'; erasing 'this is a test' turns 0 blackmail attempts into 13 — 新智元 · 2026-09-20