Pretraining recipes are vibes-based while RL is brutally mathematical
willcb · x · 2026-10-03
willcb revisits the "optimal teacher" question he pondered in April, saying the best answer he's seen so far comes from a thread on training methodology.
His key observations:
- Pretraining recipes are extremely vibes-based: industry practice is largely empirical — "we upsampled this website because the vibes were good, and we reformatted the metadata."
- RL is brutally mathematical: e.g., "we traced the error accumulation of low precision kernels and inverted some PDEs to obliterate all persistent bias."
He finds this middle ground between vibes-driven recipes and rigorous math a refreshing development in how models are trained.
Related event: Researcher: Pretraining Relies on Intuition While RL Is Hard Math(2 posts)→
More from Models
- Rumored GPT-6 'Bel' from OpenAI to face off against Anthropic's Fable 5.5 lineup — CurieuxExplorer · 2026-10-03
- V4.1 trains smoothly without dense warmup tuning: more complex algorithm, simpler information flow — teortaxesTex · 2026-10-03
- Fable 5.5 hallucinates less and replicates small models like Jev in hours, claims Bindu Reddy — bindureddy · 2026-10-03
- Ten Days of Nonstop Releases: Gemini, GPT, Sonnet and More New Models — haider1 · 2026-10-03
- Opus 5.5 and Sonnet 5.5 Now Available in Google Antigravity — NBMVegeta · 2026-10-03
- llama.cpp PR Halves Indexer Score Memory for Qwen Flash, Cutting VRAM Use — jacek2023 · 2026-10-03