~90% of frontier lab compute now goes to post-training and inference
IanAndrewsDC · x · 2026-10-06
A widely shared datapoint claims roughly 90% of frontier lab compute is now allocated to post-training and inference rather than pretraining — a sign the industry's compute重心 has shifted toward RL fine-tuning and serving real inference workloads.
More from Infra
- Can a 128GB M5 Max Mac Studio Handle Concurrent Local LLM Agents? — Simple_Telephone_867 · 2026-10-06
- MiniMax discloses 70-80% inference margins, fueling debate on AI subscription subsidies — menhguin · 2026-10-06
- KLIF open-sources a Rust front-end that manages llama.cpp, vLLM and TTS servers in one window — Koksny · 2026-10-06
- Qwen 27B on 2× RX 7900 XT: 66.5 TPS single-stream, still short of claimed 100+ — EqualCryptographer67 · 2026-10-06
- Dev burns 842B tokens in September — $409k at API list price, pays just 3.4% via subscription — doodlestein · 2026-10-06
- One Dot burns 1.6B tokens/day on a $100 subscription — roughly $540k/month in API-equivalent compute — DarthSilent · 2026-10-06