Tiny neural nets hit emergent basic reasoning at 6,616 tok/s on a single Intel AMX core
GregoryDiamos · x · 2026-09-08
Greg Diamos set out to solve an engineering problem, not write a paper: with no GPU cluster budget, he needed a model doing simple extraction and classification at 10k tokens/sec per CPU core. He gave Claude Code a pile of tokens to build one, and got three unexpected discoveries plus a paper and model.
Key constraints and findings:
- Every run is pinned to one physical core (Intel Xeon Silver 4514Y, OMPNUMTHREADS=1) — a roofline exercise, not a stunt
- AMX is worth 9x over AVX-512 fp32 at the model's shapes: best single bf16 GEMM at 2,231 GF/s, bf16 at 1024³ reaching 976–1,074 GF/s vs 121 GF/s for fp32 without AMX
- The punchline: if active parameters are small enough that hot weights fit in the 2 MiB L2 per core, 10k+ tokens/sec generation on one core is just arithmetic
The original post drew 120K views and 87 replies in a day; the writeup covers what the community pushed back on and where critics were right.
More from Infra
- PyTorch Conference China kicks off in Shanghai with keynote on open frontier AI — PyTorch · 2026-09-08
- Arm's CSS for Mobile 2 Skips the NPU, Bets On-Device AI on CPU and GPU Engines — ryanshrout · 2026-09-08
- fal extends 75% off H3 Max endpoints to Sept 15, adds real-time 1080p video — isidentical · 2026-09-08
- CPU Shortage Reaches Software Teams as AI Chip Supply Constraints Spread Beyond AI — brada · 2026-09-08
- Big Tech scouts Argentina's Patagonia for new mega data centers — Polymarket · 2026-09-08
- TimesFM 3.0 merges native MLX backend: 641 series/sec on M4 Max, no PyTorch needed — rachittshah · 2026-09-08