Expert Questions LLM Benchmarks: Manipulating Metrics for Desired Conclusions
felix_red_panda · x · 2026-08-03
Commenting on recent LLM benchmark results, Felix pointed out that this is a prime example of how you can reach any conclusion if you strictly control what you benchmark.
He specifically criticized the tok/sec/$ metric for batch size 1 as being genuinely hilarious, noting the absurdity of claiming not to care about costs while only opting for the second-best option in the comparison.
Related event: Expert Questions LLM Benchmarks: Cherrypicking Leads to Desired Conclusions(2 posts)→
More from Models
- Qwen3.8-Max Launch Sparks Debate: China Leads in Positive AI Vision — joao_gante · 2026-08-04
- OpenAI's Upcoming 'Astra' Model to Focus on Multi-Agent Collaboration — thesaraharminta · 2026-08-04
- Testing Qwen3.8-Max: Open Models Are Catching Up with Closed Frontier — dair_ai · 2026-08-04
- Kimi K3 Estimated at 2.8T Params, Potentially Distilled from Smaller Opus — gabriberton · 2026-08-04
- Expert Replicates OpenAI Astra Math Results in 24 Hours — Gary Marcus · 2026-08-04
- Claude 3.5 Sonnet and Opus Are Remarkably Terrible at Making Slides — nathanbenaich · 2026-08-04