Repetition Mismatch paper shows why data mixture experiments don't scale at EMNLP
MarekRei · x · 2026-09-05
- The paper "Repetition Mismatch: Why Data Mixture Experiments Don't Scale and How to Fix Them" has been accepted to EMNLP2026 in Budapest.
- It investigates how dataset repetitions affect optimal data mixing in data-constrained training settings, showing that mixture experiment findings break down at scale due to repeated data, and proposes fixes.
- The author shared a thread with details.
More from Research
- YC-backed MovingAtomsLab banned from DeepMind's Physics-IQ Verified benchmark for 3 months — HildeKuehne · 2026-09-06
- Is a schema-aware memory graph 'overfitting'? Dev asks for the cleanest leakage test — chaachans · 2026-09-06
- Burkov: 2026 is putting recurrence back into the Transformer it removed in 2017 — burkov · 2026-09-06
- MIT study: 83% of ChatGPT essay writers couldn't quote a single line they just wrote — victor_explore · 2026-09-06
- Carbon nanocone + fullerene check valve shows >10,000x rectification in MD sims — jwt0625 · 2026-09-06
- Formalize All Human Math in a Year? Bold AI Plan Gets Eric Weinstein's Backing — AccBalanced · 2026-09-06