CPU-only inference on a $100 Celeron SBC shows 0.6B models are usable

tre7744 · reddit · 2026-07-24

CPU-only inference on a Celeron N5095 SBC shows that 0.6B models are usable, while 8B is not.

A Reddit user benchmarked six models on a Youyeetoo X1S single-board x86 machine with a Celeron N5095, 16GB RAM, and a 128GB NVMe, using Ollama in CPU-only mode. The board costs roughly $100–$130 depending on configuration.

Key results:

The author concludes that very small models can work well for classification, routing, and summarization on low-power x86 boxes, but 8B models hit a memory-bandwidth wall. Next up is a CPU vs Vulkan comparison using llama.cpp on the integrated GPU.

Original post →

More from Infra

Infra channel →