Matthew Berman: Claiming a model runs on a machine is meaningless without stating tokens/s
MatthewBerman · x · 2026-10-12
Matthew Berman argues that saying which model you can run on a machine is meaningless without stating throughput in tokens per second. Usability hinges on inference speed, and the same model can vary several-fold across hardware — a common omission in local/edge deployment claims.
More from Infra
- Is local AI trending toward GPU-interconnect-friendly workloads? — Dathide · 2026-10-12
- Project Maya runs GLM-5.3-Flash (321B MoE) locally at up to 118 tok/s on 4×4090s — inthesearchof · 2026-10-12
- What's the best local coding setup for 16GB VRAM right now? — ECrispy · 2026-10-12
- Usage dashboard shows 1,100+ cloud VMs spun up, one account with 738 machines — aniketmaurya · 2026-10-12
- Local AI on AMD Strix Halo Writes Full Tech Specs: 5-10x Slower but It Works — julianharris · 2026-10-12
- jax-graft: an AI-built JAX backend runs JAX on Apple Silicon GPUs — twiecki · 2026-10-12