Tiny Neural Nets Revival: Transformer Hits ~1500 tok/s on M4 CPU via SME2 Instructions
GregoryDiamos · x · 2026-09-14
Developer Gregory Diamos revisits 'outrageously small' neural nets: needing a 10k tok/s CPU model for data processing, he had Claude Code build one, yielding three interesting findings. Inspired by that work, another developer adapted the approach to a transformer using the SME2 matrix instructions on a MacBook Air M4, hitting roughly 1500 tok/s with several modifications. The thread highlights both an overlooked tiny-model CPU inference path and using coding agents to produce first-hand engineering results.
More from Infra
- Used RTX 5090 listed at £3,900 (~$5,200), more than double its MSRP — julianharris · 2026-09-14
- Why Amazon and Microsoft Are Taking Communities' Side Against Utilities — pstAsiatech · 2026-09-14
- Redditor builds dual AI workstation with 2x RTX 3090 and 4x Tesla P100, asks how to optimize — FearFactory2904 · 2026-09-14
- agi-memory: SQLite-only persistent memory MCP server for coding assistants, 32MB RAM — Rude_Gate7599 · 2026-09-14
- $3000 home server with 128GB VRAM runs Qwen3.8-next at 1.3k tps prefill, 70 tps code — Thin_Pollution8843 · 2026-09-14
- Top 10 foundry revenue hits $53.5B in Q2, TSMC holds 72.5% share — Beth_Kindig · 2026-09-14