SemiAnalysis: pretraining compute share to collapse from 67% to 7% as RL and inference surge

FinanceYF5 · x · 2026-09-17

SemiAnalysis estimates OpenAI and Anthropic's compute mix is flipping: pretraining will drop from 67% to 7% of total compute, while post-training/RL rises from 5% to 55% and inference reaches 38%.

The piece also charts HBM trends: next-gen accelerators shift from 12-hi to 8-hi stacks as standard — Nvidia's Rubin Ultra drops to 192GB from 288GB per GPU — reversing earlier expectations of 16-hi and beyond. For bandwidth-sensitive inference, 4-hi HBM delivers the best $/bandwidth and lowest cost per token; beyond a capacity threshold, extra HBM yields diminishing returns at constant BOM cost. Major labs' ASIC teams plan to adopt this from HBM4 onward. Memory bandwidth, not just compute, is the next battleground.

Related event: Pretraining compute share to plummet from 67% to 7%: SemiAnalysis(2 posts)→

Original post →

More from Infra

Infra channel →