LLM Leaderboards Manipulated by Eval Config: Same Model Scores 31% to 89%
rohanpaul_ai · x · 2026-09-01
A study shows LLM leaderboard results are highly sensitive to evaluation configuration. Keeping models and 3,679 questions fixed, changing only prompt format, option order, and scoring method caused gemma4-31b to score between 31% and 89%, and 4 of 12 models reached rank 1 under some configuration. 95.7% of average gaps between neighboring models come from config-fragile items. The biggest instability source is scoring method: generating answers vs. choosing highest-likelihood option.
More from Models
- Z.ai Releases GLM-5.3-Flash: 320B Params, 1M Context, and NVFP4 Quantization — alejandroll10 · 2026-09-01
- Open Source Models Shift to Revenue Sharing and Licensing — zephyr_z9 · 2026-09-01
- Has anyone tuned a model to operate exclusively in E-prime yet? — cephaloform · 2026-09-01
- Heavy users report Claude quality dropping over the past week: eager to execute, no more clarifying questions — Numerous_Leopard_522 · 2026-09-01
- User seeks best LLM for CLI coding on single 3080 Ti — -samae1- · 2026-09-01
- Lan Hackathon Review: Qwen Excels in Physics/Engineering, K3 in General Intelligence — 葬AI · 2026-09-01