Dev's CPU-native LLM Architecture Hits 113-130 tok/s on a 10B Model, Quality Lags

WildPino25 · reddit · 2026-09-15

Reddit user WildPino25 presents SiliconLLM, a CPU-native LLM architecture that runs a 10B-parameter model at 113-130 tok/s on a Ryzen 5 3600X with no GPU — though the weights are low quality. To test quality at scale, he started training a 206M model on a T4 (8+ weeks) and tried converting Qwen2.5-Coder to his SSM/ternary/sparse format, finding donor adaptation painfully hard. The repo and research branch are on GitHub.

Original post →

More from Infra

Infra channel →