Dual 3090s Run 75B Model: Inference Optimization Tested

_ballzdeep_ · reddit · 2026-07-16

This post shares a complete hands-on test and configuration of running NVIDIA Nemotron-Labs-3-Puzzle-75B-A9B on 2x RTX 3090 GPUs. The focus isn't on introducing the model, but on optimizing the local inference stack.

Key takeaways include:

The author shared the following benchmark results:

Original post →

More from Infra

Infra channel →