Speculation: OpenAI's efficient models free compute for bigger ones; Opus 5.5 uses 4x output tokens
haider1 · x · 2026-09-25
The author speculates OpenAI's efficient gpt-6 sol/luna models were chosen to free data center capacity for a larger unreleased model (bigger than astra). Notably, Opus 5.5 uses 4x more output tokens — if API pricing roughly reflects model size, that could mean up to 8x more compute at those benchmark scores. Unverified speculation.
More from Infra
- Engineer argues HBM stacks dilute bandwidth to 20% of a single DRAM layer — knowrohit07 · 2026-09-25
- Caching policy lookups per inode cuts eBPF security agent CPU cost ~90% — JeremyCMorgan · 2026-09-25
- Photonic matrix core on thin-film lithium niobate runs in-situ backpropagation at 8-bit precision — jwt0625 · 2026-09-25
- Report: Google to Launch TPUs Into Space Next Week Aboard Falcon 9 to Test Orbital AI Data Centers — VariationLivid3193 · 2026-09-25
- Opus 5.5 Tops SimpleBench at 88.4%, Xiaomi Open-Sources MiMo-V2.6-Pro Near Frontier — Latent Space · 2026-09-25
- Skip the PC case: open-air rigs are the way for 3+ RTX Pro 6000 setups — knowrohit07 · 2026-09-25