Blind arena update: flash-next beats 3.5 max while Gemini 3.1 pro lags behind
Tall_Abrocoma_3533 · reddit · 2026-09-08
Another update to the anonymized (blind) model arena, this time with frontier models tested. Highlights from the author:
- flash-next outperforming 3.5 max
- Gemini 3.1 pro ranking surprisingly far behind
Blind rankings tend to surface gaps in human preference that official benchmarks miss; the full board is in the linked post.
More from Models
- Users slam Anthropic's Astra as slow and verbose despite frontier-level capability — jarrodwatts · 2026-09-09
- GPT-6 Astra's output reportedly reduced a user to tears, outclassing Fable — teortaxesTex · 2026-09-09
- Magic's roadmap: long-context RL, latent-knowledge alignment, then a model release — magicailabs · 2026-09-09
- Thomson Reuters' frontier-competitive legal model Thomson trained for just $450K with curated data — schwarzjn_ · 2026-09-09
- Scoop: Anthropic reportedly prepping fresh Fable-class pretrain to launch before September IPO — teortaxesTex · 2026-09-09
- User claims GPT-6 'Astra' is a step-function leap in generality, effectively AGI — brandon_galang · 2026-09-09