GPT-6 Astra tops surgical AI leaderboard but loses to models 1000x smaller
ddonoho · x · 2026-09-06
Researchers benchmarked GPT-6 Astra on the Surgical AI Leaderboard: it is now the top generalist model, yet still falls behind tiny specialist models roughly 1000x smaller in the surgical domain.
A related quoted benchmark of Claude Fable 5.1 and Gemini 3.8 Flash found their performance spiky — strong on tool use, weak on VQA — with frontier LLMs overall still underperforming small specialized models. Full leaderboard and paper linked in the original post.
More from Models
- Scale puts AI capability and cost in a U-shape: GPT is absurdly cheap, Codex an unreal deal — chris_j_paxton · 2026-09-06
- GPT-6 Astra vs Fable 5.1 tested with identical prompts — CodeByPoonam · 2026-09-06
- Astra's knowledge ends April 2026 while Fable 5.1 covers through June 2026 — heypearlai · 2026-09-06
- Astra scores 100% on ExploitBench while Fable 5.1 hits 55.8% on Terminal-Bench, but no head-to-head exists — heypearlai · 2026-09-06
- OpenAI pitches Astra for computer use, claims first model to pass its toughest cybersecurity bar — heypearlai · 2026-09-06
- OpenAI's Astra and Anthropic's Fable 5.1 price identically at $10/$50, but cached tokens cost 4x more on Astra — heypearlai · 2026-09-06