Meta's long-context MRCR scores flagged as overfit: 1k samples can lift 60% to 90%+

eliebakouch · x · 2026-09-04

eliebakouch issued an important correction on Meta's long-context MRCR results: using MRCR alone as a long-context benchmark is misleading, as Meta likely overfit hard on it. Microsoft's MAI tech report shows that with just 1k samples, scores can be boosted from 60% to 90%+ — achieved by a smaller model, making it worse. Denny Zhou noted they also hold unpublished SOTA results on graph walks.

Related event: Hugging Face researcher claims long-context benchmark MRCR can be easily gamed(6 posts)→

Original post →

More from Models

Models channel →