MRCR likely benchmaxed: 1k samples boost long-context scores from 60% to 90%+

eliebakouch · x · 2026-09-03

Researcher eliebakouch argues the long-context benchmark MRCR is likely heavily overfit: Microsoft's MAI tech report shows just 1k synthetic samples can lift scores from 60% to 90%+ — achieved by an even smaller model, making it worse. Muse Spark 1.3's strong MRCR doesn't carry over to other long-context benchmarks like graphwalks, and most vendors reporting MRCR probably include synthetic examples in training. Microsoft says there's no evidence this yields general gains, and Anthropic stopped publishing MRCR months ago. He calls for diverse benchmarks and test-time scaling curves instead of raw scores.

Related event: MRCR Long-Context Benchmark Sparks Overfitting Debate(2 posts)→

Original post →

More from Models

Models channel →