MRCR Long-Context Benchmark Sparks Overfitting Debate

A researcher claimed the MRCR long-context benchmark can be gamed with just 1k samples, lifting scores from 60% to 90%. Google author Kiran Vodrahalli responded that MRCR measures sequential understanding rather than retrieval, fueling ongoing debate.

2026-09-03 ~ 2026-09-03 · 2 related posts