Ollama beats llama.cpp by 16-35% on Celeron N5095 CPU; Vulkan cuts temps 27C but breaks on big models

tre7744 · reddit · 2026-09-02

A rigorous matched benchmark of Ollama 0.32.1 vs llama.cpp (commit 9a286ac) on a low-power Celeron N5095 board, using identical prompts, Q4KM files, 4,096-token context, four threads, and 96 generated tokens.

Findings

Bottom line: keep Ollama for CPU, restrict Vulkan to Qwen3 0.6B/1.7B, leave MTP off. Full write-up, scripts and failure timeline are public.

Original post →

More from Infra

Infra channel →