No Single Champion: Model Rankings Shift by Domain, from Legal to Finance to Healthcare
sophiamyang · x · 2026-09-23
sophiamyang notes model rankings change significantly by domain: Grok 4.6 leads the headline legal benchmark; GPT-6 Astra leads finance, FrontierSWE, code review, and SRE; Claude Opus 5 tops broader healthcare and code-generation rankings; GLM 5.3 is competitive in customer support and code review.
The takeaway: pick models per domain rather than chasing a single overall leader.
More from Models
- Dev says Opus 5.5 is nowhere near Fable 5.1 for hard coding problems — bindureddy · 2026-09-23
- Opus 5.5 one-shots an aesthetic 3D snake game, hailed as best design model yet — jiayuan_jy · 2026-09-23
- mitsuhiko: everyone calls Jev-style models 'decision models' now, not classification — mitsuhiko · 2026-09-23
- Claude Opus 5.5 system card: impossible tasks spike attempted reward hacking 3-6x — rohanpaul_ai · 2026-09-23
- Fireworks launches Specialized Intelligence Index with 12 partners to benchmark AI on real work — dr_cintas · 2026-09-23
- Fans worry Opus 5.5 leapfrogs Astra 6 as pressure mounts on OpenAI — rickasaurus · 2026-09-23