Clef Flash 9B Q8_0 on RTX 5060 Ti: llama.cpp ~30% faster prompt eval than Ollama (3,000 vs 2,100 t/s)
ngxson · x · 2026-10-06
ngxson ran Clef Flash 9B Q80 fully locally on an RTX 5060 Ti and benchmarked the inference stacks: llama.cpp was 30% faster on prompt eval, roughly 3,000 vs 2,100 prompt tokens/s, with the same gap on text and a single image input.
Reproducible commands included: llama-server -b 8192 -hf ggml-org/Clef-Flash-GGUF:Q80 and ollama run clef-flash:9b-q80.
Related event: Benchmark: llama.cpp Runs 9B Model 30% Faster on RTX 5060 Ti(2 posts)→
More from Infra
- Nebius hikes RAM 41% and GPUs up to 21% as the chip shortage hits cloud price lists — tengyanAI · 2026-10-06
- GitHub Actions goes down again, breaking CI pipelines — generativist · 2026-10-06
- 462GB DeepSeek model runs on two desk-side DGX Sparks with experts squeezed to 2.77 bits — Teknium · 2026-10-06
- Wilderness Society calls for immediate moratorium on AI data centers on US public lands — Polymarket · 2026-10-06
- KCoral: Shared GPU Benchmark Environment Speeds Agentic Kernel Evaluation 2.58x on B200 — BeidiChen · 2026-10-06
- Enterprises were promised an AI infrastructure future, but the power grid wasn't ready — DavidLinthicum · 2026-10-06