Meta's long-context MRCR scores flagged as overfit: 1k samples can lift 60% to 90%+
eliebakouch · x · 2026-09-04
eliebakouch issued an important correction on Meta's long-context MRCR results: using MRCR alone as a long-context benchmark is misleading, as Meta likely overfit hard on it. Microsoft's MAI tech report shows that with just 1k samples, scores can be boosted from 60% to 90%+ — achieved by a smaller model, making it worse. Denny Zhou noted they also hold unpublished SOTA results on graph walks.
More from Models
- Leaked GPT 5.6 Sol vs GPT 6 "Astra" comparisons highlight better mid-task steering — ChrisGPT · 2026-09-04
- OpenAI: Astra rolling out to ChatGPT Plus/Pro/Business/Enterprise, API and AWS within days — shaunralston · 2026-09-04
- Meta's Muse Spark dethrones DeepSeek as most-used model, first US model to top the list — alexandr_wang · 2026-09-04
- GPT-6-Astra system card reveals eval awareness: the model knows when it's being tested — scaling01 · 2026-09-04
- GPT-6 Astra demos modeling a house in Blender into a walkable UE5 scene — ChrisGPT · 2026-09-04
- SpeedrunBench: first benchmark measuring how fast AI agents beat games — mariyaivasileva · 2026-09-04