llama.cpp beats Ollama by ~30% prompt eval speed on RTX 5060 Ti with Clef Flash 9B
ngxson · x · 2026-10-06
Developer ngxson benchmarked Clef Flash 9B (Q80) fully locally on an RTX 5060 Ti and found llama.cpp 30% faster at prompt eval than Ollama: roughly 3,000 vs 2,100 prompt tokens/s, with the same gap on text and single-image input.
He shared copy-paste commands for both stacks (llama-server -b 8192 -hf ggml-org/Clef-Flash-GGUF:Q80, ollama run clef-flash:9b-q80) and pointed to ggml-org's Decision models collection on Hugging Face, which includes Clef, OpenJev, Laya and other GGUF models plus small classification models.
Related event: Benchmark: llama.cpp Runs 9B Model 30% Faster on RTX 5060 Ti(2 posts)→
More from Infra
- Nebius hikes RAM 41% and GPUs up to 21% as the chip shortage hits cloud price lists — tengyanAI · 2026-10-06
- GitHub Actions goes down again, breaking CI pipelines — generativist · 2026-10-06
- 462GB DeepSeek model runs on two desk-side DGX Sparks with experts squeezed to 2.77 bits — Teknium · 2026-10-06
- Wilderness Society calls for immediate moratorium on AI data centers on US public lands — Polymarket · 2026-10-06
- KCoral: Shared GPU Benchmark Environment Speeds Agentic Kernel Evaluation 2.58x on B200 — BeidiChen · 2026-10-06
- Enterprises were promised an AI infrastructure future, but the power grid wasn't ready — DavidLinthicum · 2026-10-06