Exllamav3 benchmarks show major speedup over Llama.cpp on dual 3060s

Ecstatic-Wash-7667 · reddit · 2026-08-17

User benchmarked Exllamav3 against Llama.cpp on a dual RTX 3060 setup using Qwen models (3.8 27b and 3.6 35b) with 3 cold-start runs. Results show Exllamav3 significantly outperforms Llama.cpp in tokens/sec across tested configurations.

Original post →

More from Infra

Infra channel →