Reddit Asks: Is exl3/tabbyapi Underrated? Better Compression and Speed Than Other Backends
C1oover · reddit · 2026-08-17
Reddit user C1oover asks why exl3/tabbyapi doesn't get more attention. After testing multiple backends, they find exl3 has better compression per bit and runs Qwen 27 and Gemma 31 faster on a 3090 than optimized llama.cpp and vLLM. They wonder if they're missing something or if it's just forgotten on the subreddit.
More from Infra
- Merge Gateway adds Grok 4.6 with temporary discount — shensi · 2026-08-17
- Every dev commit now triggers massive compute amplification compared to 20 years ago — andreisavu · 2026-08-17
- Investigation reveals potential Microsoft AI chip shortage — nordicinst · 2026-08-17
- AI chips will replace crypto as the superior energy-to-value transducer — beffjezos · 2026-08-17
- Exllamav3 benchmarks show major speedup over Llama.cpp on dual 3060s — Ecstatic-Wash-7667 · 2026-08-17
- OpenSSH 10.5 Released: AI-Driven Security Reports Surge, Forcing Faster Release Cadence — chrisrohlf · 2026-08-17