NInfer Enables Efficient Qwen3.8-27B Inference on RTX 3090
mrmontanasagrada · reddit · 2026-08-15
A developer ported the NInfer runtime to RTX 3090 (24GB) to optimize Qwen3.8-27B inference. Benchmarks show the model achieves 71 tokens/s for single requests and 165 tokens/s decoding across 8 concurrent requests, utilizing ReplaySSM and CUDA Graphs while staying within 24GB VRAM. The project demonstrates the potential of running large models on older consumer GPUs, with preliminary results also mentioned for the Qwen3-35B-A3B model.
More from Infra
- Tips for debugging Training-Inference Mismatch (T.I.M) — rosstaylor90 · 2026-08-15
- Soleio: Compute is sovereignty — soleio · 2026-08-15
- CEO fixes MoE kernel index overflow bug in sonic-moe — rosstaylor90 · 2026-08-15
- Long Reasoning Task Consumes 22k Tokens for 3k Output — generativist · 2026-08-15
- NVIDIA Local AI systems now support more open-source models — nvidia · 2026-08-15
- Q-CTRL quantum setup runs materials sim 3,000x faster than classical solvers — MJBiercuk · 2026-08-15