Qwen3.8-27B Inference Speed: ~30-32 t/s on RTX 3090

CooLittleFonzies · reddit · 2026-08-17

A Reddit user shares their local inference speed for Qwen3.8-27B: with an RTX 3090, 64GB DDR4, and AMD 7950x, using Q5KM GGUF quantization, they get about 30-32 tokens/s. The post aims to gather other users' hardware and speed data for comparison.

Original post →

More from Models

Models channel →