Open-ended RL is 'orbit by freefall': only progress per GPU-time matters

tensorqt · x · 2026-09-17

Discussion around Periodic's results: Neon was midtrained and RL'd on proprietary lab data for the XRD domain, yet general-purpose Astra lands within 2pp on FrontierXRD at 40% higher cost—near-specialist performance without specialist training. tensorqt adds the deeper takeaway: for open-ended tasks, as in pretraining, all that matters is progress per GPU-time—general models that keep improving will close the gap.

Original post →

More from Models

Models channel →