Surgical VLM Leaderboard: All Frontier Models Fall Far Short of Specialized Models
ddonoho · x · 2026-09-15
The SDSC × UChicago team updated its Surgical Intelligence Leaderboard, benchmarking 20+ vision-language models across 10 surgical datasets covering instrument recognition, anatomy, skill assessment, context/VQA, and recommendations.
Key findings:
- Every general-purpose VLM is significantly below the specialized-model ceiling (set at 1.0). The best, GPT-6 Astra, scores 0.566; Qwen3.8 Max hits 0.423; Gemini models cluster around 0.37-0.40; Claude/Grok/Kimi mostly fall between 0.1-0.35, with some small models scoring near or below zero.
- A fine-tuned small model, LemonFM (linear probe), reaches 0.863 — far above all frontier models and near specialist level.
- Per-modality breakdown shows skill assessment is nearly unsolved for general models, while recommendation tasks are comparatively stronger.
Takeaway: on surgical video understanding, the gap between general VLMs and domain specialists remains enormous.
More from Models
- Resemble AI ships DETECT-World, a physics-based deepfake detector with 99.5% audio accuracy — AiBreakfast · 2026-09-15
- Grok 4.7 Misses Target Again; 2.5T-Parameter Grok 4.8 Finishes Training — eyishazyer · 2026-09-15
- Writers Report Gemini Flash 3.8 Suffers Severe Long-Context Rot Despite 1M Token Window — Quenty1 · 2026-09-15
- Ai2 Chair's 3 Predictions: Open-Source AI Goes from Ideology to Infrastructure in 18 Months — billhilf · 2026-09-15
- Why Surgery Benchmarks Reward Fine-tuned Small Models While Math Benchmarks Don't — ddonoho · 2026-09-15
- Claude Opus 5.2 spotted in grayscale testing on Claude Code, seemingly skipping 5.1 — Angaisb_ · 2026-09-15