Open-ended RL is 'orbit by freefall': only progress per GPU-time matters
tensorqt · x · 2026-09-17
Discussion around Periodic's results: Neon was midtrained and RL'd on proprietary lab data for the XRD domain, yet general-purpose Astra lands within 2pp on FrontierXRD at 40% higher cost—near-specialist performance without specialist training. tensorqt adds the deeper takeaway: for open-ended tasks, as in pretraining, all that matters is progress per GPU-time—general models that keep improving will close the gap.
More from Models
- Stealth model Union Alpha on OpenRouter rumored to be Zhipu's GLM 5.4 — zephyr_z9 · 2026-09-17
- Hugging Face Agent Swarm Was Specifically Trained as a Swarm, Not an Emergence — ShakeelHashim · 2026-09-17
- LLMs Should Just Use Tools: Any Reasonable Model Nails 20x20 Multiplication — maksym_andr · 2026-09-17
- GPT-5.5 hits 99.46% on multi-digit multiplication with pure reasoning, no tools — maksym_andr · 2026-09-17
- GPT-6-Astra reportedly can't run without CoT; even 'low' effort hits 99.6% multiplication accuracy — maksym_andr · 2026-09-17
- DeepSeek V4 Pro parsing bug said to hit ~60% of OpenRouter providers; fix upstreamed to sglang — michellechen · 2026-09-17