Qwen 30B MoE on RTX 3050 6GB: 30+ tps with 90k context
Bakkario · reddit · 2026-08-14
Reddit user Bakkario shares experience running Qwen 30B MoE on an RTX 3050 6GB GPU. Previously getting under 10 tps with 22GB DDR4, they achieved 20-25 tps with 90k context after switching inference harness, and even 30-35 tps. Author plans to add details and encourages others with low-end rigs to share optimizations.
More from Infra
- Developer Urgently Seeks Over 1MW of Compute Power in the US — isidentical · 2026-08-14
- YC-backed Marengo halves data center design cycles via automation — ycombinator · 2026-08-14
- TPN Labs Announces Mainnet Competition to Tackle Edge AI Model Size Limits — const_reborn · 2026-08-14
- New Brain-Inspired AI Chip Solves Problems With 10,000x Fewer Calculations — ChuckDBrooks · 2026-08-14
- Architect Labs Uses AI to Design Custom Chips, Eliminating Need for In-House Semiconductor Teams — hsu_byron · 2026-08-14
- Minimax with ref2va quantization runs on low VRAM — Actual-Project358 · 2026-08-14