Expert Questions LLM Benchmarks: Manipulating Metrics for Desired Conclusions

felix_red_panda · x · 2026-08-03

Commenting on recent LLM benchmark results, Felix pointed out that this is a prime example of how you can reach any conclusion if you strictly control what you benchmark.

He specifically criticized the tok/sec/$ metric for batch size 1 as being genuinely hilarious, noting the absurdity of claiming not to care about costs while only opting for the second-best option in the comparison.

Related event: Expert Questions LLM Benchmarks: Cherrypicking Leads to Desired Conclusions(2 posts)→

Original post →

More from Models

Models channel →