Dual 3090s Run 75B Model: Inference Optimization Tested
_ballzdeep_ · reddit · 2026-07-16
This post shares a complete hands-on test and configuration of running NVIDIA Nemotron-Labs-3-Puzzle-75B-A9B on 2x RTX 3090 GPUs. The focus isn't on introducing the model, but on optimizing the local inference stack.
Key takeaways include:
- By using a specific quantized version, it can run a full 262K context on dual 3090s without any CPU offload.
- A crucial breakthrough was tweaking PYTORCHCUDAALLOCCONF to resolve fragmentation and rank imbalance issues.
- PIECEWISE CUDA graphs provided the biggest decode boost: jumping from roughly 27.9 tok/s to 88+ tok/s.
- Enabling custom all-reduce further increased throughput for N=1 and N=4.
The author shared the following benchmark results:
- N=1 decode at 93.8 tok/s, prefill at roughly 3660 tok/s
- N=4 aggregate at 255 tok/s, TTFT p50 at 0.61s, with no preempts or OOM errors
- Passed the 52K needle recall test, and tool calling functions normally
More from Infra
- Why a 1GW Chinese AI data center may be plausible after all — teortaxesTex · 2026-07-22
- China’s AI arms race is increasingly defined by chips, data centers, and open models — BenBajarin · 2026-07-22
- Agent search bottlenecks are now about variance, not raw latency — rohanpaul_ai · 2026-07-22
- Gavin Baker argues Nvidia may be one of open source AI’s biggest supporters — GavinSBaker · 2026-07-22
- AI Power Demand Exposes US Energy Gap, Urging Shift from Scarcity to Abundance — bradneuberg · 2026-07-22
- Gavin Baker says Nvidia’s $630B figure would be system revenue, not all Nvidia’s — GavinSBaker · 2026-07-22