AI Benchmarks Questioned: Meta and Others Accused of Data Manipulation

iruletheworldmo · x · 2026-07-09

The commentary argues that current AI benchmarks are becoming increasingly irrelevant, urging the public to stay critical of scores released by vendors. The post specifically calls out the head of Meta AI, claiming that models underperformed in the "Humanity's Last Exam"—which the executive helped design—compared to the coding benchmarks of GPT 5.5 and Claude Opus. Furthermore, it alleges that the core benchmarks for Grok 4.5 were entirely compromised and custom-designed by their team.

Related event: AI Benchmark Credibility Questioned Amid Meta's Results(2 posts)→

Original post →

More from Models

Models channel →