Artificial Analysis Shakes Up Rankings Again, Accused of Skipping Proper Evaluations
Tall_Abrocoma_3533 · reddit · 2026-09-08
A Redditor notes that model evaluation platform Artificial Analysis has updated yet again, significantly reshuffling rankings — but complains it did not run proper evaluations before pushing the benchmark update. The post is thin on detail, mainly reflecting community frustration with the platform's methodology and update cadence.
More from Models
- Magic's roadmap: long-context RL, latent-knowledge alignment, then a model release — magicailabs · 2026-09-09
- Thomson Reuters' frontier-competitive legal model Thomson trained for just $450K with curated data — schwarzjn_ · 2026-09-09
- User claims GPT-6 'Astra' is a step-function leap in generality, effectively AGI — brandon_galang · 2026-09-09
- OUI-1: first open-weights Generative UI model hits 71.7% with just 4B params — iamrobotbear · 2026-09-09
- How GPT-6 Astra's computer use works: it rides on the accessibility tree — iamrobotbear · 2026-09-09
- DeepSeek V4.1 'humiliates' rival in creative game design, RL approach hailed as vindicated — teortaxesTex · 2026-09-09