Frontier pre-training runs likely capped at ~2 months to avoid wasting algorithmic progress
scaling01 · x · 2026-09-07
A case that frontier pre-training runs are now capped at roughly two months: spending 100 days on 100K GB200s makes little sense, and long runs lock you out of rapid algorithmic progress — effectively wasting 100 days of it. The author takes this as the base case for GPT-6 Astra's pre-training.
More from Infra
- FP8/FP4 Quantization Delivers Only ~1.5x and 2x Real Speedups, Far Below Theoretical Gains — scaling01 · 2026-09-07
- Naura demos key etch process for 64-layer 3D DRAM without EUV, selectivity above 500:1 — pstAsiatech · 2026-09-07
- Netherlands builds 'Dutch AI' by finetuning Qwen 3.5 27B in subsidized datacenter — teortaxesTex · 2026-09-07
- This Week's AI Must-Reads: OpenAI's Research Acceleration Report and Broadcom's $16.7B AI Chip Quarter — VibeMarketer_ · 2026-09-07
- On-device Android agent with Gemma 4 E2B hits 2.6 tok/s live vs 11 tok/s on replay — HowDevelop · 2026-09-07
- Dev wrote Marlin-style FP4/FP8 kernels for RTX 3090 — NVIDIA declined to upstream them — QuixiAI · 2026-09-07