Puzzle-75B-A9B Inference Performance on Three 3090s

Important_Quote_1180 · reddit · 2026-07-09

The post detailed the NVFP4 inference performance of NVIDIA Puzzle-75B-A9B on 3x 3090 GPUs: achieving a decoding speed of approximately 132 tokens/s under vLLM 0.22.1, with a total system wall power consumption of about 500W.

Original post →

More from Infra

Infra channel →