Researcher calls for long-context detail: high MRCR scores don't track real performance

eliebakouch · x · 2026-09-03

HF researcher eliebakouch replies to Jack Rae arguing that MRCR is a very easy long-context benchmark to game — both Microsoft's MAI tech report and Anthropic note that higher MRCR scores (likely via training on synthetic data similar to the benchmark) don't correlate with real long-context performance for users.

He calls on labs to release more technical detail on long-context training, speculating others may be doing the same without disclosing it, and suggests the upcoming open-weight Muse Spark release could be an opportunity to share such details.

Related event: MRCR Long-Context Benchmark Accused of Being Benchmaxed; Author Responds(6 posts)→

Original post →

More from Models

Models channel →