NInfer Adds Day-0 Support for Qwen3.8-27B, Hits ~200 tok/s on RTX 5090

FormOne2615 · reddit · 2026-08-15

NInfer inference engine now supports Qwen3.8-27B on day zero, achieving 200 tok/s generation on a single RTX 5090 with speculative decoding.

Key improvements include:

Weights are available on Hugging Face; feedback and bug reports are welcome.

Original post →

More from Infra

Infra channel →