27B model on a single RTX 4090: 262K context at ~130 tok/s with NInfer

Distinct-Pie2389 · reddit · 2026-10-09

The author ran the uncensored HauhauCS Qwen 27B model on a single RTX 4090 using NInfer, their own C++/CUDA inference engine, after converting from GGUF:

Converter and writeup are open-sourced in the ninfer-4090 repo and PR on GitHub.

Original post →

More from Infra

Infra channel →