Epoch AI Gives Models 3,000 GPU-Hours to Invent a New Post-Training Method Beyond GRPO
rms80 · x · 2026-10-10
Epoch AI gave Fable 5 and GPT-5.6 Sol 3,000 GPU-hours each to develop a novel post-training technique improving on GRPO. markcummins calls it the only eval worth tracking: models remain poor at invention — a modest conceptual leap like GRPO to SDPO proves harder than recent math results, and when that changes, all bets are off.
More from Models
- Cloudflare releases clef-omni, an open omni-modal model with audio, image and video input — ritakozlov · 2026-10-10
- AWS Bedrock posts legacy notices for Claude Opus 4.1, Sonnet 4 and Sonnet 4.5 — repligate · 2026-10-10
- Google slammed for not releasing Argon after officially announcing it — almmaasoglu · 2026-10-10
- Meta paper shows byte-level models beat tokenized ones given enough training compute — alex_verem · 2026-10-10
- Dev warns after Anthropic terms: diversify your toolchain or ideology compliance may cost you — AlexTensor · 2026-10-10
- Step 5 Preview now available on OpenRouter — kimmonismus · 2026-10-10