Surgical VLM Leaderboard: All Frontier Models Fall Far Short of Specialized Models

ddonoho · x · 2026-09-15

The SDSC × UChicago team updated its Surgical Intelligence Leaderboard, benchmarking 20+ vision-language models across 10 surgical datasets covering instrument recognition, anatomy, skill assessment, context/VQA, and recommendations.

Key findings:

Takeaway: on surgical video understanding, the gap between general VLMs and domain specialists remains enormous.

Original post →

More from Models

Models channel →