LLM Leaderboards Manipulated by Eval Config: Same Model Scores 31% to 89%

rohanpaul_ai · x · 2026-09-01

A study shows LLM leaderboard results are highly sensitive to evaluation configuration. Keeping models and 3,679 questions fixed, changing only prompt format, option order, and scoring method caused gemma4-31b to score between 31% and 89%, and 4 of 12 models reached rank 1 under some configuration. 95.7% of average gaps between neighboring models come from config-fragile items. The biggest instability source is scoring method: generating answers vs. choosing highest-likelihood option.

Original post →

More from Models

Models channel →