llama.cpp Fork Adds Qwen3.8-Flash Support, Hits 44 tks on RTX 5090 at Q4

giveen · reddit · 2026-08-28

Developer giveen submitted PR #324 to TheTom/llama-cpp-turboquant, porting qwen4exp (Qwen3.8-Flash-Next) support from upstream vLLM PR #27742. Benchmarked on an RTX 5090 at Q4 quantization with --moe-cache auto enabled, it reaches roughly 44 tokens/s.

Original post →

More from Infra

Infra channel →