Gaperon paper shows late test-set contamination recovers benchmark scores without hurting generation
burny_tech · x · 2026-09-09
A threaded discussion around the open Gaperon paper (arXiv:2510.25771): the team released 1.5B/8B/24B French-English models trained on 2-4T tokens with full pipelines and hundreds of checkpoints. Key findings: linguistic quality filtering improves fluency but lowers benchmark scores, while deliberate late contamination with test sets recovers competitive scores with only reasonable harm to generation; neural filtering can unintentionally amplify benchmark leakage. Responder soldni flags two confounders for newer corpora scoring higher: test-set contamination and eval hacking, since benchmarks like OLMES are commonly used during data-recipe development, so newer mixtures are likely optimized for them. A methodological warning that rising benchmark scores may not reflect real capability gains.
More from Models
- OpenAI launches academic program, with a caveat: use at your own risk — lucacarlone1 · 2026-09-09
- A breakdown of the most popular GLiNER models for zero-shot NER — kalyan_kpl · 2026-09-09
- NSA/FBI/CISA advisory urges silent model degradation over suspected distillation — xuanalogue · 2026-09-09
- OpenAI's Navier–Stokes run: train-while-deploying and 10,000 coordinated agents — DataLearnerAI · 2026-09-09
- GPT-6 Astra rumored to hit ~11.6-hour p80 task horizon, tracking AI-2027 curve — haider1 · 2026-09-09
- Astra makes a weird dashboard mistake at just 44% context usage — eigenron · 2026-09-09