48-hour inference run estimated at $13.8M
GregKamradt · x · 2026-08-28
A back-of-the-napkin calculation estimates a 48-hour inference run at $13.824 million, assuming 400M tokens/min throughput and a 50/50 input/output split without prompt caching. Real-world costs would likely be lower due to caching and log optimizations.
Related event: Estimated 48-Hour Inference Cost for OpenAI Event: $13.8M(2 posts)→
More from Infra
- ASML mirror supply becomes the new AI compute bottleneck — TheZvi · 2026-08-28
- NVIDIA Details QAD Pipeline for Optimizing Nemotron Model — PyTorch · 2026-08-28
- Developer Runs 291B Model Locally on Four Mac Studios — eptwts · 2026-08-28
- Cloudflare opens monetization gateway for APIs and MCP tools — kleffew94 · 2026-08-28
- Survey: 72% of Devs Prefer 5x Speed Over 20% Smarter Models — sonic_op · 2026-08-28
- SOMA's OpenClaw compression core deployed in production — markjeffrey · 2026-08-28