llama.cpp beats Ollama by ~30% prompt eval speed on RTX 5060 Ti with Clef Flash 9B

ngxson · x · 2026-10-06

Developer ngxson benchmarked Clef Flash 9B (Q80) fully locally on an RTX 5060 Ti and found llama.cpp 30% faster at prompt eval than Ollama: roughly 3,000 vs 2,100 prompt tokens/s, with the same gap on text and single-image input.

He shared copy-paste commands for both stacks (llama-server -b 8192 -hf ggml-org/Clef-Flash-GGUF:Q80, ollama run clef-flash:9b-q80) and pointed to ggml-org's Decision models collection on Hugging Face, which includes Clef, OpenJev, Laya and other GGUF models plus small classification models.

Related event: Benchmark: llama.cpp Runs 9B Model 30% Faster on RTX 5060 Ti(2 posts)→

Original post →

More from Infra

Infra channel →