Reddit user flags Artificial Analysis as unreliable: same model shows conflicting scores

metigue · reddit · 2026-09-04

A Reddit user investigated the recent score of 61 for Astra on Artificial Analysis and found troubling inconsistencies: different pages report different scores for the same model on the same benchmark, and some figures directly contradict the benchmark providers' own verified results. The author concludes the site can no longer be trusted as a data source.

Compounding the problem, up-to-date versions of direct benchmarks like DeepSWE and terminal bench are missing many models — especially open-source ones — making alternatives scarce. The post asks the community for better model comparison resources.

Original post →

More from Models

Models channel →