Grok 4.7 vs Claude Fable 5.1 vs GPT-6 Astra vs DeepSeek V4.1 Flash: 40+ benchmark showdown
HealthySkeptic2000 · reddit · 2026-09-22
A large snapshot (Sept 21, 2026) aggregating Artificial Analysis and Vals AI results across four frontier models: - Overall: GPT-6 Astra and Claude Fable 5.1 tie on AA Intelligence Index (53); Fable 5.1 leads Vals Index (68.83%); Grok 4.7 at 46, DeepSeek V4.1 Flash at 39. - Coding: GPT-6 Astra tops Terminal-Bench 4.0 (59% AA), IOI (100%), and code migration; Fable 5.1 leads Vibe Code Bench (90.26%). - Knowledge/reasoning: Fable 5.1 best on HLE (59%, 65% with tools) but with a 73% hallucination rate; DeepSeek worst at 96%. - Professional work: Fable 5.1 dominates legal/finance/Excel benchmarks; Grok 4.7 only leads Harvey Legal Agent (19.58%). Overall Fable 5.1 and GPT-6 Astra take most category wins; Grok and DeepSeek lag notably.
More from Models
- Xiaomi open-sources MiMo-V2.6 omni-modal models, topping open-model index at 46.32 — victormustar · 2026-09-22
- SemiAnalysis says open source is dying, yet 20+ open models shipped in the past month — _lewtun · 2026-09-22
- Grok 4.7 posts 59% recall on defensive cyber bench at half the cost of rivals — andreamichi · 2026-09-22
- Grok 4.7 jumps from #9 to #3 on BuildingBench with 0.783, 66% cheaper than Fable 5.1 — ZhitingHu · 2026-09-22
- Jev explained: why the AI community's new favorite isn't a traditional LLM — multiply_matrix · 2026-09-22
- LLMs excel at 1-token output — dev proposes replacing low/medium/high reasoning tiers with token counts — arkuto · 2026-09-22