Baseten's head of model training explains why RL extends reasoning but generalization needs domain-specific post-training

baseten · x · 2026-09-23

Connor O'Neill, Baseten's Head of Model Training, joined Dwarkesh Patel's podcast to break down horizon generalization: RL teaches models to work longer, but reasoning remains dependent on domain-specific post-training. The episode also covers what's next at the frontier of training longer-horizon models.

Original post →

More from Models

Models channel →