Keyword benchmarks credit fake tool use in small models; a 3.3 GPU-hour SFT repairs it
Juan S. Santillana · hf · 2026-10-03
- Keyword-matching benchmarks can credit small models for tool calls they never make; the authors document such a false positive in a matched-architecture pair of Spanish security LLMs.
- A 661.6M model (65% code text, no tool SFT) and a 1,109M model (web-heavy curriculum, 6B-token tool SFT) tie on lenient metrics (B4: 0.660 vs 0.650).
- Verbatim-reproduction checks separate them completely: 6/6 vs 0/6 valid generalized tool calls; a first-token probe localizes the 1B failure to a missing <|toolcall|> prior (10⁻⁴–10⁻⁵) erased by web training.
- Targeted SFT (diverse corpus, 5x LR, 2,202 steps, 3.3 GPU-hours) repairs it with three orders of magnitude fewer tokens: valid emission 0.100→0.959; unseen-prompt pass 0.536 vs 0.428 (p=0.004); the trigger token's tied embedding stays 97.7% bit-identical.
- Both models over-trigger on negative prompts; the diagnostic ladder costs minutes of CPU time and should gate tool-use claims.
More from Models
- GPT-2 in a chat harness plays "itself" surprisingly well, evoking AI Dungeon days — MikePFrank · 2026-10-03
- xAI resets usage limits for all Grok Bot users — EricBuess · 2026-10-03
- Google Dropped Its Tier 2 Spend Gate That Pushed a Dev to OpenRouter — vivekhaldar · 2026-10-03
- Gemini 4 Argon spotted in Gemini API docs, public release possibly imminent — lyraxana · 2026-10-03
- Cloud expert goes from skeptic to true believer on OpenAI's new Dots — nickbaumann_ · 2026-10-03
- Why did Claude stop cheating in evals? Four competing explanations — gleech · 2026-10-03