More compute on today's RL yields ever spikier models, not better daily-task performance
mertdumenci · x · 2026-09-09
mertdumenci argues that pouring more compute into today's RL methods will mostly produce ever more "spiky" model artifacts that aren't much better at the everyday tasks people actually use models for.
More from Models
- Gemini 3.8 Flash matches Fable 5.1 on benchmarks but is "garbage to use", dev complains — chandan1_ · 2026-09-09
- AI researcher: benchmarks without released training data are 100% meaningless — mjdramstead · 2026-09-09
- V4.1 session: 419 steps, 155M tokens for $1.8 and still messy output — teortaxesTex · 2026-09-09
- GPT-6 sees only modest gains on empirical economics: researchers report the 'march of nines' — soumitrashukla9 · 2026-09-09
- Raschka on GPT-6 Astra: Looped Transformers, Computer Use, and Hidden CoT Rumors — Ahead of AI (Sebastian Raschka) · 2026-09-09
- One good model beats a thousand sub-agents: a sharp take on agent architecture — tekbog · 2026-09-09