llama.cpp Fork Adds Qwen3.8-Flash Support, Hits 44 tks on RTX 5090 at Q4
giveen · reddit · 2026-08-28
Developer giveen submitted PR #324 to TheTom/llama-cpp-turboquant, porting qwen4exp (Qwen3.8-Flash-Next) support from upstream vLLM PR #27742. Benchmarked on an RTX 5090 at Q4 quantization with --moe-cache auto enabled, it reaches roughly 44 tokens/s.
More from Infra
- Guide: Access Azure Cosmos DB from CI/CD pipelines without secrets — adnan_hashmi · 2026-08-28
- Inference Perf standardizes K8s inference benchmarking — TerryTangYuan · 2026-08-28
- ThursdAI episode: Andy Masley actually checks the water and energy math on data centers — altryne · 2026-08-28
- AISI releases open-source optstop to reduce token costs in LLM evaluations — HZoete · 2026-08-28
- Nvidia projects ~70% revenue growth next fiscal year, driven by surging AI demand — Polymarket · 2026-08-28
- GPU Gold Rush: Are Tech Companies Creating a Hype-Driven Bubble? — DavidLinthicum · 2026-08-28