Dev's CPU-native LLM Architecture Hits 113-130 tok/s on a 10B Model, Quality Lags
WildPino25 · reddit · 2026-09-15
Reddit user WildPino25 presents SiliconLLM, a CPU-native LLM architecture that runs a 10B-parameter model at 113-130 tok/s on a Ryzen 5 3600X with no GPU — though the weights are low quality. To test quality at scale, he started training a 206M model on a T4 (8+ weeks) and tried converting Qwen2.5-Coder to his SSM/ternary/sparse format, finding donor adaptation painfully hard. The repo and research branch are on GitHub.
More from Infra
- Helion closes oversubscribed $500M Series G for fusion energy — ycombinator · 2026-09-16
- nginx 1.31.6 Patches Heap Buffer Overflow in HTTP/3 (CVE-2026-90439) — jedisct1 · 2026-09-16
- AWS Trainium Runs PyTorch Natively, PyTorch Con Keynote Revealed — PyTorch · 2026-09-16
- A long-form explainer on why local inference matters — and why you need uncensored models — HankYeomans · 2026-09-15
- DGX Spark + TRELLIS.2 turns pencil sketches into 3D GLBs fully locally — jasonkneen · 2026-09-15
- Skipping a second RTX 5090 for two more Sparks: local inference user explains why — ideamaker321 · 2026-09-15