Sample selection and ordering matter a lot in LLM training: DataFlex makes data scheduling dynamic
Puzzleheaded_Box2842 · reddit · 2026-09-11
The author argues that as training data scales, simple shuffling is no longer enough: noisy/redundant samples, domain imbalance, and example ordering all shape what LLMs learn. They outline four dynamic scheduling levers—dynamic selection (using loss, gradient similarity, or offline scores to pick samples per training window), dynamic reordering for curriculum-style training, dynamic mixing of domain proportions, and dynamic weighting of gradient contributions. The core idea: treat data scheduling as part of optimization rather than a fixed pre-training decision. This is implemented in the open-source OpenDCAI/DataFlex project.
More from Research
- 3D ResNet Paper Crosses 3,000 Citations Eight Years After CVPR 2018 — HirokatuKataoka · 2026-09-11
- Jeff Heaton's Intro to the Math of Neural Networks eBook Is Free to Download — blaizedsouza · 2026-09-11
- Mathematician Daniel Litt Launches Problem Repo to Track Human vs AI Progress: 15 Problems, 1 Solved — littmath · 2026-09-11
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Navier-Stokes, Riemann, P vs NP: what this week's math buzzwords mean for you — koltregaskes · 2026-09-11