Reddit user flags Artificial Analysis as unreliable: same model shows conflicting scores
metigue · reddit · 2026-09-04
A Reddit user investigated the recent score of 61 for Astra on Artificial Analysis and found troubling inconsistencies: different pages report different scores for the same model on the same benchmark, and some figures directly contradict the benchmark providers' own verified results. The author concludes the site can no longer be trusted as a data source.
Compounding the problem, up-to-date versions of direct benchmarks like DeepSWE and terminal bench are missing many models — especially open-source ones — making alternatives scarce. The post asks the community for better model comparison resources.
More from Models
- Critics warn OpenAI's GPT-6 Astra reasons opaquely, gutting CoT monitoring safety — GaryMarcus · 2026-09-04
- ChatGPT adds writing-style matching from connected apps, analytics, and a Yubikey deal tied to Daybreak access — btibor91 · 2026-09-04
- GPT-6 reportedly launches as Tesla starts public rides in steering-free Cybercab — Dr_Singularity · 2026-09-04
- Astra early-access users' similar blender demos look coordinated, with no practical examples shown — jdjohnson · 2026-09-04
- Researcher teases dynamic composite eval index as "evals run on Twitter vibes" — evijit · 2026-09-04
- OpenAI engineer says GPT-6 "Astra" has "tremendously" improved writing quality — mark_k · 2026-09-04