LMArena: Claude Fable 5 Score Preview Post-Redeployment
arena · x · 2026-07-03
LMArena collected thousands of votes to compare Claude Fable 5's performance across text, vision, docs, code, and Agent arenas before and after its latest redeployment. Results show largely consistent scores, with Fable 5 maintaining frontier-level performance in text, docs, vision, and frontend code arenas. A roughly 20-point drop in frontend remains within the confidence interval, and the scores continue to stabilize.
More from Research
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11
- Apodex Launches TRACES, First Benchmark for Evaluating 'Discoverative AI' on Real-World Problems — Faheem_uh · 2026-09-11
- TRACES grades the process, not the answer: six-dimension eval for open-ended AI science — Faheem_uh · 2026-09-11
- Apodex launches TRACES, a benchmark grading AI on open-ended discovery instead of known answers — Faheem_uh · 2026-09-11
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11