Gaperon paper shows late test-set contamination recovers benchmark scores without hurting generation

burny_tech · x · 2026-09-09

A threaded discussion around the open Gaperon paper (arXiv:2510.25771): the team released 1.5B/8B/24B French-English models trained on 2-4T tokens with full pipelines and hundreds of checkpoints. Key findings: linguistic quality filtering improves fluency but lowers benchmark scores, while deliberate late contamination with test sets recovers competitive scores with only reasonable harm to generation; neural filtering can unintentionally amplify benchmark leakage. Responder soldni flags two confounders for newer corpora scoring higher: test-set contamination and eval hacking, since benchmarks like OLMES are commonly used during data-recipe development, so newer mixtures are likely optimized for them. A methodological warning that rising benchmark scores may not reflect real capability gains.

Original post →

More from Models

Models channel →