Why 20-hour, 200M-token tasks make mass post-training nearly impossible

nrehiew_ · x · 2026-09-03

Long-horizon tasks serve as a proxy for out-of-distribution evaluation: with single tasks taking 20 hours and 200M tokens, mass post-training on them is extremely difficult — an infra nightmare involving compaction and more — which is precisely what makes them a meaningful test of generalization.

Related event: FrontierSWE v2 Launches: 20-Hour Ultra-Long-Horizon Coding Benchmark, Fable 5.1 Leads by Over 24 Points(5 posts)→

Original post →

More from Infra

Infra channel →