Why 20-hour, 200M-token tasks make mass post-training nearly impossible
nrehiew_ · x · 2026-09-03
Long-horizon tasks serve as a proxy for out-of-distribution evaluation: with single tasks taking 20 hours and 200M tokens, mass post-training on them is extremely difficult — an infra nightmare involving compaction and more — which is precisely what makes them a meaningful test of generalization.
More from Infra
- Broadcom guides AI revenue to ~$115B in FY2027, doubling again to $230B in FY2028 — BenBajarin · 2026-09-03
- Perplexity open-sources Lily, its Apple Silicon inference engine for Qwen3.6-35B-A3B — inductionheads · 2026-09-03
- Nvidia now ~8% of S&P 500 market cap, worth 16.3% of US GDP at $5.3T — ivan_bezdomny · 2026-09-03
- Broadcom Q3 AI chip revenue hits $16.7B, up 221% YoY; guides $21.7B for Q4 — Beth_Kindig · 2026-09-03
- PINNACLE: Claude Fable 5.1 halves agent failure rate but costs $2.46 per correct answer — ryanshrout · 2026-09-03
- eBPF looks cheap but isn't free: Bitbison's deep dive on hooks, CO-RE reads and rings — tianyin_xu · 2026-09-03