MRCR likely benchmaxed: 1k samples boost long-context scores from 60% to 90%+
eliebakouch · x · 2026-09-03
Researcher eliebakouch argues the long-context benchmark MRCR is likely heavily overfit: Microsoft's MAI tech report shows just 1k synthetic samples can lift scores from 60% to 90%+ — achieved by an even smaller model, making it worse. Muse Spark 1.3's strong MRCR doesn't carry over to other long-context benchmarks like graphwalks, and most vendors reporting MRCR probably include synthetic examples in training. Microsoft says there's no evidence this yields general gains, and Anthropic stopped publishing MRCR months ago. He calls for diverse benchmarks and test-time scaling curves instead of raw scores.
Related event: MRCR Long-Context Benchmark Sparks Overfitting Debate(2 posts)→
More from Models
- Muse Spark 1.3 ships with stronger instruction following and long-horizon agentic coding — armand_ruiz · 2026-09-03
- Fully automated tracker discovers new Hugging Face models within a day and estimates max tok/s — helloiamleonie · 2026-09-03
- Astra rumored to drop today as labs stagger releases knowing rivals' plans — haider1 · 2026-09-03
- Crazy week of model drops; open-weight Chinese copies seen arriving in 3-5 months — scottleibrand · 2026-09-03
- Sam Altman jokes GPT-6 will be renamed GPT-6-7 — Miles_Brundage · 2026-09-03
- Sitemap sleuthing points to 1pm PT today as OpenAI's likely Astra launch window — imjustnewatai · 2026-09-03