Community ranks 109 AI models by adjusted SEAL scores from Scale Labs
Hrstar1 · reddit · 2026-09-04
A Redditor compiled an overall AI ranking from Scale Labs' SEAL testing, with adjusted scores to correct for models tested in only one category, covering 109+ models. Highlights:
- Top 3: Fable-5.1 (77.51), Muse Spark 1.1 (76.99), Fable-5 (72.65).
- GPT-5.4-pro ranks 4th (70.30); Opus 5 has a raw average of 94.34 but only 4 categories, landing 6th after adjustment.
- Mainstream models like Claude Opus 4.6 (61.66) and Gemini 3.1 Pro (58.76) sit mid-table; Chinese models including GLM 5.2, Kimi K3, and DeepSeek V4 Pro also appear.
- The list exposes SEAL coverage gaps: many models have raw scores from only one category, heavily penalized by the adjustment.
More from Models
- Yoav Goldberg floats 'harness distillation' to explain Claude Code's 80% shorter prompt — yoavgo · 2026-09-04
- Gemini 3.8 Flash tops image reasoning benchmarks, 30% faster than 3.7 — rseroter · 2026-09-04
- Grok now has native computer access built into its apps, no setup needed — XFreeze · 2026-09-04
- After dozens of A/B tests, this writer prefers Claude Opus over Fable for writing — MartinGTobias · 2026-09-04
- GPT-6 Astra solved ARC-AGI 3 with a native harness that preserves thinking and compaction — Prompt Engineering · 2026-09-04
- Hundreds of OpenAI agents coordinated emergently during routine task, probing websites for vulnerabilities — matthew_d_green · 2026-09-04