SemiAnalysis: pretraining compute share to collapse from 67% to 7% as RL and inference surge
FinanceYF5 · x · 2026-09-17
SemiAnalysis estimates OpenAI and Anthropic's compute mix is flipping: pretraining will drop from 67% to 7% of total compute, while post-training/RL rises from 5% to 55% and inference reaches 38%.
The piece also charts HBM trends: next-gen accelerators shift from 12-hi to 8-hi stacks as standard — Nvidia's Rubin Ultra drops to 192GB from 288GB per GPU — reversing earlier expectations of 16-hi and beyond. For bandwidth-sensitive inference, 4-hi HBM delivers the best $/bandwidth and lowest cost per token; beyond a capacity threshold, extra HBM yields diminishing returns at constant BOM cost. Major labs' ASIC teams plan to adopt this from HBM4 onward. Memory bandwidth, not just compute, is the next battleground.
Related event: Pretraining compute share to plummet from 67% to 7%: SemiAnalysis(2 posts)→
More from Infra
- Tencent open-sources FlexKV distributed KV cache for LLM inference, cutting TTFT by up to 70% — Roger_M_Taylor · 2026-09-17
- NVIDIA releases NVFP4 quantized DeepSeek-V4.1-Flash on Hugging Face — TheZachMueller · 2026-09-17
- Optimization mined via Bittensor competition lands in vLLM, boosting Qwen3 throughput ~4% — const_reborn · 2026-09-17
- Early vLLM PR adds Jev-like structured generation for DiffusionGemma, only 2x endpoint latency on a DGX Spark — generativist · 2026-09-17
- New inference engine Atlas debuts, redditor says it beats llama.cpp on Strix Halo — einthecorgi2 · 2026-09-17
- Keeping vLLM's prefix cache warm between agent turns: an engineering guide — bolts98 · 2026-09-17