Running a 26B Model on an Old Xeon

neomindryan · hn · 2026-07-15

This article demonstrates running Gemma 4 26B on a 13-year-old Xeon machine without a GPU, achieving a speed of about 5 tokens/s.

The focus is on viable inference on low-end hardware. The author uses this post to illustrate that running large models doesn't strictly require modern GPUs.

Original post →

More from Infra

Infra channel →