Pretraining recipes are vibes-based while RL is brutally mathematical

willcb · x · 2026-10-03

willcb revisits the "optimal teacher" question he pondered in April, saying the best answer he's seen so far comes from a thread on training methodology.

His key observations:

He finds this middle ground between vibes-driven recipes and rigorous math a refreshing development in how models are trained.

Related event: Researcher: Pretraining Relies on Intuition While RL Is Hard Math(2 posts)→

Original post →

More from Models

Models channel →