Exllamav3 benchmarks show major speedup over Llama.cpp on dual 3060s
Ecstatic-Wash-7667 · reddit · 2026-08-17
User benchmarked Exllamav3 against Llama.cpp on a dual RTX 3060 setup using Qwen models (3.8 27b and 3.6 35b) with 3 cold-start runs. Results show Exllamav3 significantly outperforms Llama.cpp in tokens/sec across tested configurations.
More from Infra
- Stripe acquired OpenRouter for $7 billion — sven_ai · 2026-08-17
- AI Spending Bubble: Hyperscaler Commitments Top $3 Trillion — GaryMarcus · 2026-08-17
- AI's massive energy footprint: Pathways to net positive impact — AryHHAry · 2026-08-17
- Merge Gateway adds Grok 4.6 with temporary discount — shensi · 2026-08-17
- Every dev commit now triggers massive compute amplification compared to 20 years ago — andreisavu · 2026-08-17
- Investigation reveals potential Microsoft AI chip shortage — nordicinst · 2026-08-17