MRCR Long-Context Benchmark Sparks Overfitting Debate
A researcher claimed the MRCR long-context benchmark can be gamed with just 1k samples, lifting scores from 60% to 90%. Google author Kiran Vodrahalli responded that MRCR measures sequential understanding rather than retrieval, fueling ongoing debate.
2026-09-03 ~ 2026-09-03 · 2 related posts
- MRCR likely benchmaxed: 1k samples boost long-context scores from 60% to 90%+ — eliebakouch · 2026-09-03
- MRCR author pushes back on benchmaxing claims: it measures ordering, not needle retrieval — eliebakouch · 2026-09-03