Grok 4.7 vs Claude Fable 5.1 vs GPT-6 Astra vs DeepSeek V4.1 Flash: 40+ benchmark showdown

HealthySkeptic2000 · reddit · 2026-09-22

A large snapshot (Sept 21, 2026) aggregating Artificial Analysis and Vals AI results across four frontier models: - Overall: GPT-6 Astra and Claude Fable 5.1 tie on AA Intelligence Index (53); Fable 5.1 leads Vals Index (68.83%); Grok 4.7 at 46, DeepSeek V4.1 Flash at 39. - Coding: GPT-6 Astra tops Terminal-Bench 4.0 (59% AA), IOI (100%), and code migration; Fable 5.1 leads Vibe Code Bench (90.26%). - Knowledge/reasoning: Fable 5.1 best on HLE (59%, 65% with tools) but with a 73% hallucination rate; DeepSeek worst at 96%. - Professional work: Fable 5.1 dominates legal/finance/Excel benchmarks; Grok 4.7 only leads Harvey Legal Agent (19.58%). Overall Fable 5.1 and GPT-6 Astra take most category wins; Grok and DeepSeek lag notably.

Original post →

More from Models

Models channel →