Tiny Neural Nets Revival: Transformer Hits ~1500 tok/s on M4 CPU via SME2 Instructions

GregoryDiamos · x · 2026-09-14

Developer Gregory Diamos revisits 'outrageously small' neural nets: needing a 10k tok/s CPU model for data processing, he had Claude Code build one, yielding three interesting findings. Inspired by that work, another developer adapted the approach to a transformer using the SME2 matrix instructions on a MacBook Air M4, hitting roughly 1500 tok/s with several modifications. The thread highlights both an overlooked tiny-model CPU inference path and using coding agents to produce first-hand engineering results.

Original post →

More from Infra

Infra channel →