NInfer Enables Efficient Qwen3.8-27B Inference on RTX 3090

mrmontanasagrada · reddit · 2026-08-15

A developer ported the NInfer runtime to RTX 3090 (24GB) to optimize Qwen3.8-27B inference. Benchmarks show the model achieves 71 tokens/s for single requests and 165 tokens/s decoding across 8 concurrent requests, utilizing ReplaySSM and CUDA Graphs while staying within 24GB VRAM. The project demonstrates the potential of running large models on older consumer GPUs, with preliminary results also mentioned for the Qwen3-35B-A3B model.

Original post →

More from Infra

Infra channel →