GPT-6 Astra tops surgical AI leaderboard but loses to models 1000x smaller

ddonoho · x · 2026-09-06

Researchers benchmarked GPT-6 Astra on the Surgical AI Leaderboard: it is now the top generalist model, yet still falls behind tiny specialist models roughly 1000x smaller in the surgical domain.

A related quoted benchmark of Claude Fable 5.1 and Gemini 3.8 Flash found their performance spiky — strong on tool use, weak on VQA — with frontier LLMs overall still underperforming small specialized models. Full leaderboard and paper linked in the original post.

Original post →

More from Models

Models channel →