Epoch AI Gives Models 3,000 GPU-Hours to Invent a New Post-Training Method Beyond GRPO

rms80 · x · 2026-10-10

Epoch AI gave Fable 5 and GPT-5.6 Sol 3,000 GPU-hours each to develop a novel post-training technique improving on GRPO. markcummins calls it the only eval worth tracking: models remain poor at invention — a modest conceptual leap like GRPO to SDPO proves harder than recent math results, and when that changes, all bets are off.

Original post →

More from Models

Models channel →