Pretraining data mixing is 'vibes-based' while RL is brutally mathematical, says willcb

willcb · x · 2026-10-03

AI researcher willcb contrasts the two halves of LLM training: pretraining data recipes are famously vibes-based ('we upsampled this site because the vibes were good, also reformatted metadata'), while RL is brutally mathematical (tracing error accumulation in low-precision kernels, inverting PDEs to eliminate persistent bias). He argues the sweet spot is using inference compute to optimally interpolate between teacher and student — 'it's just MCMC' — and that this exposes why pure teacher SFT and naive OPSD are fundamentally flawed.

Related event: Researcher: Pretraining Relies on Intuition While RL Is Hard Math(2 posts)→

Original post →

More from Research

Research channel →