AMD5 Local Inference Hits Less Than 0.2 tok/s
DanGrover · x · 2026-07-13
Testing on a mid-range machine with an AMD5 chip and 128GB of memory yielded an inference speed of less than 0.2 tok/s. Even enabling features like offloading some experts to the GPU didn't noticeably improve the speed. The author concludes that while it's very slow, the fact that it "runs at all" is still pretty cool.
Related event: GLM CPU-only Inference Questioned as Speed Falls Below 0.2 tok/s(2 posts)→
More from Infra
- Local AI may pay back in 6–7 years and cut long-term costs by 30–40% — DavidLinthicum · 2026-07-21
- TSMC reportedly plans up to 10% chipmaking price hikes in 2027 — kimmonismus · 2026-07-21
- More open models and llama.cpp updates are coming, says Merve Noyan — mervenoyann · 2026-07-21
- Why adding a second LLM provider breaks more than the API surface — Ok_Extension6373 · 2026-07-21
- UK AI datacentres face backlash over heat, noise and land use — nordicinst · 2026-07-21
- Fluidstack raises $830M at $7.5B valuation as Anthropic backs a $50B compute buildout — rohanpaul_ai · 2026-07-21