Qwen 30B MoE on RTX 3050 6GB: 30+ tps with 90k context

Bakkario · reddit · 2026-08-14

Reddit user Bakkario shares experience running Qwen 30B MoE on an RTX 3050 6GB GPU. Previously getting under 10 tps with 22GB DDR4, they achieved 20-25 tps with 90k context after switching inference harness, and even 30-35 tps. Author plans to add details and encourages others with low-end rigs to share optimizations.

Original post →

More from Infra

Infra channel →