Active params in L2 cache hit 14,572 tok/s on Intel Emerald Rapids with AMX
GregoryDiamos · x · 2026-09-07
Gregory Diamos needed a 10k tok/s CPU model for data processing, so he gave Claude Code a pile of tokens to build an outrageously small neural net. With active params fitted into L2 cache on an Intel Emerald Rapids CPU with AMX, it reaches 14,572 tok/s. He shares three interesting discoveries and argues tiny nets deserve a revisit for inference/data-processing workloads.
Related event: Developer Uses Claude Code to Build Tiny CPU Model Hitting 10k+ tok/s(4 posts)→
More from Infra
- Dual RTX 5080 and 128GB RAM: What Models to Run on 32GB VRAM — whatyathinkk · 2026-09-07
- CXMT's response to Apple partnership rumor sounds like a tacit confirmation — zephyr_z9 · 2026-09-07
- AMD bets on Personal AI at IFA 2026: 192GB unified memory platform and a 96-core liquid-cooled desktop supercomputer — APPSO · 2026-09-07
- Lenovo Yoga Pro 9n runs a 120B-parameter model locally in a 1.65kg Windows laptop, with 128GB unified memory — APPSO · 2026-09-07
- Malaysia weighs Huawei 910C for $494M sovereign AI project despite US warnings — kimmonismus · 2026-09-07
- CubeSandbox v0.7.0 ships cross-node pause & resume for AI agent sandboxes — HeyAmit_ · 2026-09-07