Toby Ord: RL's scaling surprise may hinge on mid-training, benchmark gains may overstate progress
tobyordoxford · x · 2026-09-24
Toby Ord discusses RL scaling: generalization to non-verifiable tasks is better than expected, possibly due to mid-training; or models may just be learning what's tested, making benchmarks misrepresent overall progress.
More from Models
- TeleOCR, a Qwen2.5-VL-based document parsing model, trends on Hugging Face — StarDoc-AI · 2026-09-24
- Rumor: SSI to launch its first model this month after security-related delay — iruletheworldmo · 2026-09-24
- Which sub-40B finetunes work best for mimicking a writing style? — Borkato · 2026-09-24
- First-day Opus 5.5 verdict: power user says it replaced Astra entirely — kimmonismus · 2026-09-24
- Opus 5.5 impresses early users; Mirage launch video made with just 4 turns of edits — seanwbren · 2026-09-24
- Open source multimodal decision model XOR launches on Hugging Face, 260k context, Qwen-based — TheMoonMidas · 2026-09-24