What data are labs using to train rumored 10T-parameter models?
Ill_Fisherman8352 · reddit · 2026-07-25
The post asks what data mix labs are using to train rumored 10T-parameter models.
It frames the problem as a scaling mismatch: if parameter counts are rising 3x, does training data need to rise proportionally too? The author points to the long-running “data wall” concern and asks whether the extra data is coming from:
- model-generated reasoning chains during inference
- human-generated reasoning traces collected via services like Mercor
- some other synthetic or filtered data mix
The post is essentially a request for the current state of large-model data sourcing and how labs are getting enough tokens at this scale.
More from Research
- A 10-question LLM regime-reasoning test probes ambiguity handling and falsifiability — Local-Reading-1624 · 2026-07-25
- An essay links compression and intelligence to mark Ray Solomonoff’s 100th birthday — ryangr · 2026-07-25
- 10 agent eval patterns every AI engineer should know, from golden sets to trajectory scoring — Roger_M_Taylor · 2026-07-25
- A copy-paste SOP aims to make LLM analysis more reliable — Local-Reading-1624 · 2026-07-25
- HUG uses 1M egocentric frames to train zero-shot robot grasping — chris_j_paxton · 2026-07-25
- NeurIPS paper proposes CAPA to show similar models may weaken AI oversight — dhadfieldmenell · 2026-07-25