Serving 2.8T Param Models Gets Cheap: $1.75/M Tokens on AMD MI355X
scaling01 · x · 2026-07-31
A developer provided perspective on the plummeting serving costs for massive models. According to napkin math and test compiles, Luminal could potentially hit $1.75 / M output tokens on AMD MI355X and $2.92 / M tokens on B300 in the coming months. While still a work in progress, it highlights a crazy world of rapidly improving inference economics for 2.8T parameter models.
More from Infra
- AWS Revenue Grows 37%, Annual Sales for Custom AI Chips Exceed $25B — 智东西 · 2026-07-31
- Revisiting Lossy Verification in Speculative Decoding: Mechanisms and Failure Modes — Tianyu Wang · 2026-07-31
- Energy Consumption: Single AI Prompt vs Agentic Workflow Differs by 100,000x — AndyMasley · 2026-07-31
- UBS Chart Highlights the Central Hub of the AI Compute Supply Chain — BenBajarin · 2026-07-31
- Stacking 512GB VRAM: Developer Builds Dual-Node 8x V100 Inference Cluster — UltraFOV · 2026-07-31
- Open-source Rust GGUF runtime runNburn runs 295B model on 64GB RAM, 2.8x faster decode than llama.cpp — coderyeon · 2026-07-31