AMD5 Local Inference Hits Less Than 0.2 tok/s
DanGrover · x · 2026-07-13
Testing on a mid-range machine with an AMD5 chip and 128GB of memory yielded an inference speed of less than 0.2 tok/s.
Even enabling features like offloading some experts to the GPU didn't noticeably improve the speed. The author concludes that while it's very slow, the fact that it "runs at all" is still pretty cool.
Related event: GLM CPU-only Inference Questioned as Speed Falls Below 0.2 tok/s(2 posts)→
More from Infra
- 12 KV Cache Reduction Techniques Every AI Engineer Should Understand, Explained — blaizedsouza · 2026-09-11
- The shadow GPU capacity market is formalizing, with Meta selling excess compute to outside buyers — DavidLinthicum · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- 80% of the DIY LLM inference hype posters have already quit — it's brutally hard systems work — abhijithneil · 2026-09-11
- Hugging Face's Ultra Scale Playbook: a free book on training LLMs on GPU clusters — mdancho84 · 2026-09-11
- Is inference latency becoming the biggest bottleneck for production AI agents? — Euphoric_Sea632 · 2026-09-11