Pretraining data mixing is 'vibes-based' while RL is brutally mathematical, says willcb
willcb · x · 2026-10-03
AI researcher willcb contrasts the two halves of LLM training: pretraining data recipes are famously vibes-based ('we upsampled this site because the vibes were good, also reformatted metadata'), while RL is brutally mathematical (tracing error accumulation in low-precision kernels, inverting PDEs to eliminate persistent bias). He argues the sweet spot is using inference compute to optimally interpolate between teacher and student — 'it's just MCMC' — and that this exposes why pure teacher SFT and naive OPSD are fundamentally flawed.
Related event: Researcher: Pretraining Relies on Intuition While RL Is Hard Math(2 posts)→
More from Research
- Causal Inference Assumptions: What to Check Before You Run the Model — Nasereliver · 2026-10-03
- V4.1 trains smoothly without dense warmup tuning: more complex algorithm, simpler information flow — teortaxesTex · 2026-10-03
- Teaching AI to Drive with Evolution: An Intuitive Neuroevolution Explainer — egehancry · 2026-10-03
- Why AI can't solve the mystery of time: training presupposes the very clock it must explain — johnseach · 2026-10-03
- arXiv trends suggest AI uplift hits quantum physics within a year, all physics by late 2029 — cephaloform · 2026-10-03
- Berkeley's Humanoid Intelligence Center wins both tracks at IROS 2026 RoCo Challenge — berkeley_ai · 2026-10-03