Why Surgery Benchmarks Reward Fine-tuned Small Models While Math Benchmarks Don't
ddonoho · x · 2026-09-15
The author shares a non-obvious insight from their surgery benchmark: specialized-skill benchmarks don't all behave like math benchmarks.
- On math benchmarks (e.g., FrontierMath), frontier models dominate in descending order of size/release date, and fine-tuned mini-models are non-players.
- On the surgery benchmark, however, fine-tuned small models (like LemonFM) top the leaderboard — and the benchmarks were created by independent teams, so there's no self-grading bias.
- The proposed explanation: essentially all of math is "publishable" and thus embedded in pretraining data, while surgical practice knowledge is held by a small group and never systematically digitized, so general models can't learn it and fine-tuning becomes the only path.
Implication: the more封闭 the domain knowledge, the larger the advantage of fine-tuned small models over general frontier models.
More from Models
- Resemble AI ships DETECT-World, a physics-based deepfake detector with 99.5% audio accuracy — AiBreakfast · 2026-09-15
- Grok 4.7 Misses Target Again; 2.5T-Parameter Grok 4.8 Finishes Training — eyishazyer · 2026-09-15
- Writers Report Gemini Flash 3.8 Suffers Severe Long-Context Rot Despite 1M Token Window — Quenty1 · 2026-09-15
- Ai2 Chair's 3 Predictions: Open-Source AI Goes from Ideology to Infrastructure in 18 Months — billhilf · 2026-09-15
- Surgical VLM Leaderboard: All Frontier Models Fall Far Short of Specialized Models — ddonoho · 2026-09-15
- Claude Opus 5.2 spotted in grayscale testing on Claude Code, seemingly skipping 5.1 — Angaisb_ · 2026-09-15