Benchmark Manipulation: You Can Reach Any Conclusion by Controlling Tests

felix_red_panda · x · 2026-08-03

Developer @felixredpanda sharply criticized a recent model evaluation. He pointed out that by carefully selecting and controlling the specific items in a benchmark, one can reach any desired conclusion. Furthermore, he mocked the evaluation's token throughput per dollar for batch size 1 as being genuinely illogical.

Related event: Expert Questions LLM Benchmarks: Cherrypicking Leads to Desired Conclusions(2 posts)→

Original post →

More from Models

Models channel →