Researcher calls for long-context detail: high MRCR scores don't track real performance
eliebakouch · x · 2026-09-03
HF researcher eliebakouch replies to Jack Rae arguing that MRCR is a very easy long-context benchmark to game — both Microsoft's MAI tech report and Anthropic note that higher MRCR scores (likely via training on synthetic data similar to the benchmark) don't correlate with real long-context performance for users.
He calls on labs to release more technical detail on long-context training, speculating others may be doing the same without disclosing it, and suggests the upcoming open-weight Muse Spark release could be an opportunity to share such details.
Related event: MRCR Long-Context Benchmark Accused of Being Benchmaxed; Author Responds(6 posts)→
More from Models
- Prompting Fable 5.1 with embodied mannerisms like *shrugs* makes roleplay flow better — repligate · 2026-09-05
- Astra appears in ChatGPT/Codex on Windows but not Mac, same account — cheesecakegood · 2026-09-05
- No usage reset on GPT-6 Astra launch day, developer calls out OpenAI — tobowers · 2026-09-05
- GPT-6 Astra Launch Marred by Outages, OpenAI Hands Out Daily Token Reset Vouchers — 新智元 · 2026-09-05
- Debate: Models Fuzzily Recall Concepts, Not Text — SAE Features vs Edit-Distance Memorization — voooooogel · 2026-09-05
- Blogger feeds GPT6 Astra a PPT template, gets 30 conference slides with the right avatar — vista8 · 2026-09-05