14,572 tok/s on Intel Xeon: fitting active params into L2 cache with AMX
GregoryDiamos · x · 2026-09-08
The author reports experiments running LLM inference on an Intel Emerald Rapids Xeon CPU with AMX: by fitting the active parameters into the L2 cache, throughput reaches 14,572 tok/s with basic bf16. They note that better weight compression could raise the active parameter ratio further, suggesting room for optimization.
Related event: Engineer Uses Claude Code to Design Tiny Model Hitting 14572 tok/s on CPU(7 posts)→
More from Infra
- Nvidia NVL72 rack shipments forecast to grow over 50% YoY in 2027 — Beth_Kindig · 2026-09-08
- Jensen Huang confirms GPT-6 Astra trained on 100K+ Grace Blackwell NVL72 systems — rohanpaul_ai · 2026-09-08
- South Korea to give everyone free generative AI, backed by up to 512 B200 GPUs — IgorCarron · 2026-09-08
- Dev runs SDXL fine-tune fully on iPhone Neural Engine: 6-bit, 8 steps, offline — NovaDevCodeStudio · 2026-09-08
- Qwen 27B q8 vs bf16 on a DGX Spark: is the 1% token difference worth the memory? — superSmitty9999 · 2026-09-08
- Running Qwen3.8-27B for coding on 32GB VRAM — what local LLMs do you use and why? — theexile1337 · 2026-09-08