Matthew Berman: Claiming a model runs on a machine is meaningless without stating tokens/s

MatthewBerman · x · 2026-10-12

Matthew Berman argues that saying which model you can run on a machine is meaningless without stating throughput in tokens per second. Usability hinges on inference speed, and the same model can vary several-fold across hardware — a common omission in local/edge deployment claims.

Original post →

More from Infra

Infra channel →