RTX 5090 Demo: Running Qwen 3.8 27B Locally at 115 tok/s

rohanpaul_ai · x · 2026-08-17

Rohan Paul shared a demo of running Qwen 3.8 27B locally on a system with an RTX 5090 (32GB VRAM), achieving a speed of 115 tokens/sec. Notably, the model's official BF16 checkpoint is 55.6 GB, suggesting the demo likely utilized a quantized version to fit within the 32GB memory limit.

Related event: Qwen3.8-27B Benchmarked Across GPUs: Local Deployment Proves Viable(7 posts)→

Original post →

More from Embodied

Embodied channel →