Researcher: Pretraining Relies on Intuition While RL Is Hard Math
AI researcher willcb observes that pretraining data recipes rely heavily on intuition and experience, while RL post-training is grounded in hard mathematics, with newer methods sitting in between, sparking discussion on training methodology.
2026-10-03 ~ 2026-10-03 · 2 related posts
- Pretraining data mixing is 'vibes-based' while RL is brutally mathematical, says willcb — willcb · 2026-10-03
- Pretraining recipes are vibes-based while RL is brutally mathematical — willcb · 2026-10-03