Run Qwen3 at 50 tok/s Locally on RTX 5060 Ti

Azazelionide · reddit · 2026-07-11

The author conducted local inference experiments on an RTX 5060 Ti (16GB VRAM), using custom CUDA and C++ code to push Qwen3-30B-A3B to 50–54 tok/s (float8).

Original post →

More from Infra

Infra channel →