Clef Flash 9B Q8_0 on RTX 5060 Ti: llama.cpp ~30% faster prompt eval than Ollama (3,000 vs 2,100 t/s)

ngxson · x · 2026-10-06

ngxson ran Clef Flash 9B Q80 fully locally on an RTX 5060 Ti and benchmarked the inference stacks: llama.cpp was 30% faster on prompt eval, roughly 3,000 vs 2,100 prompt tokens/s, with the same gap on text and a single image input.

Reproducible commands included: llama-server -b 8192 -hf ggml-org/Clef-Flash-GGUF:Q80 and ollama run clef-flash:9b-q80.

Related event: Benchmark: llama.cpp Runs 9B Model 30% Faster on RTX 5060 Ti(2 posts)→

Original post →

More from Infra

Infra channel →