Serving 2.8T Param Models Gets Cheap: $1.75/M Tokens on AMD MI355X

scaling01 · x · 2026-07-31

A developer provided perspective on the plummeting serving costs for massive models. According to napkin math and test compiles, Luminal could potentially hit $1.75 / M output tokens on AMD MI355X and $2.92 / M tokens on B300 in the coming months. While still a work in progress, it highlights a crazy world of rapidly improving inference economics for 2.8T parameter models.

Original post →

More from Infra

Infra channel →