AI Benchmarks Questioned: Meta and Others Accused of Data Manipulation
iruletheworldmo · x · 2026-07-09
The commentary argues that current AI benchmarks are becoming increasingly irrelevant, urging the public to stay critical of scores released by vendors. The post specifically calls out the head of Meta AI, claiming that models underperformed in the "Humanity's Last Exam"—which the executive helped design—compared to the coding benchmarks of GPT 5.5 and Claude Opus. Furthermore, it alleges that the core benchmarks for Grok 4.5 were entirely compromised and custom-designed by their team.
Related event: AI Benchmark Credibility Questioned Amid Meta's Results(2 posts)→
More from Models
- A model benchmark shows Muse Spark far ahead of Grok-4.20 on score vs cost — cis_female · 2026-07-21
- A 600k-token relationship test compares how models comment on personal context — cis_female · 2026-07-21
- Cola launches July, the latest model jokingly billed as “second only to Fable” — oran_ge · 2026-07-21
- Kimi K3 looks stronger and about 5× cheaper on a frontend dashboard task — OwariDa · 2026-07-21
- Last Week in AI recap: Anthropic’s $65B round, IPO filing, and Microsoft’s MAI push — Last Week in AI · 2026-07-21
- A user says Claude 4.6 felt worse yesterday and asks whether model quality can drift over time — Rahios · 2026-07-21