27B model on a single RTX 4090: 262K context at ~130 tok/s with NInfer
Distinct-Pie2389 · reddit · 2026-10-09
The author ran the uncensored HauhauCS Qwen 27B model on a single RTX 4090 using NInfer, their own C++/CUDA inference engine, after converting from GGUF:
- Full native 262K context window
- 130 tok/s decode with MTP3 speculative decoding at 70.8% acceptance (vs 91.6 tok/s at 131K under llama.cpp)
- 3,591 tok/s prefill on a 9K prompt
- Perplexity within 1.3% of the official artifact, so the conversion is clean
- Vision and tool calls still work
Converter and writeup are open-sourced in the ninfer-4090 repo and PR on GitHub.
More from Infra
- Uber details its MCP Gateway, the unified platform all its AI agents use to reach backends — AxSaucedo · 2026-10-09
- 64GB Isn't Enough Anymore: Local AI Is Eating Through Mac RAM — gregmushen · 2026-10-09
- boat spins up 250 agent sandbox VMs for 10 cents: full Ubuntu boxes at $20/mo — RexDouglass · 2026-10-09
- A cache hit is not free: inside Triton's compilation cache and hidden costs — Mahmoud_Zalt · 2026-10-09
- AI agents could spawn history's largest bureaucracy, where machines create work for machines — brucemacv · 2026-10-09
- Bittensor-based GPU cloud Lium buys back and burns nearly $2.7M of SN51 tokens in six months — markjeffrey · 2026-10-09