14,572 tok/s on Intel Xeon: fitting active params into L2 cache with AMX

GregoryDiamos · x · 2026-09-08

The author reports experiments running LLM inference on an Intel Emerald Rapids Xeon CPU with AMX: by fitting the active parameters into the L2 cache, throughput reaches 14,572 tok/s with basic bf16. They note that better weight compression could raise the active parameter ratio further, suggesting room for optimization.

Related event: Engineer Uses Claude Code to Design Tiny Model Hitting 14572 tok/s on CPU(7 posts)→

Original post →

More from Infra

Infra channel →