NInfer Adds Day-0 Support for Qwen3.8-27B, Hits ~200 tok/s on RTX 5090
FormOne2615 · reddit · 2026-08-15
NInfer inference engine now supports Qwen3.8-27B on day zero, achieving 200 tok/s generation on a single RTX 5090 with speculative decoding.
Key improvements include:
- Up to 8 concurrent requests with a shared paged KV cache pool, each retaining full context length.
- ReplaySSM for GDN + speculative decoding, greatly reducing recurrent-state memory overhead under concurrency (not fully supported by vLLM yet).
- Numerous CUDA kernel optimizations and PDL usage to further reduce latency.
Weights are available on Hugging Face; feedback and bug reports are welcome.
More from Infra
- RTX 3090 gets 35 t/s on Qwen 3.8 27B — cviperr33 · 2026-08-15
- CME to launch futures contracts tracking Nvidia H100/B100 compute costs — AccBalanced · 2026-08-15
- Feedback: Serverless GPU capacity tight, placement times high — tobowers · 2026-08-15
- Touchmark launches futures marketplace for AI tokens, up to 30% below on-demand rates — ycombinator · 2026-08-15
- Macro Analysis: Is the AI Dip a Buy? Key Levels to Watch — Beth_Kindig · 2026-08-15
- Optimized Dual 3090 Quantization of Qwen3.8-27B Released — luedtek · 2026-08-15