Reddit Asks: Is exl3/tabbyapi Underrated? Better Compression and Speed Than Other Backends

C1oover · reddit · 2026-08-17

Reddit user C1oover asks why exl3/tabbyapi doesn't get more attention. After testing multiple backends, they find exl3 has better compression per bit and runs Qwen 27 and Gemma 31 faster on a 3090 than optimized llama.cpp and vLLM. They wonder if they're missing something or if it's just forgotten on the subreddit.

Original post →

More from Infra

Infra channel →