Qwen 3.6 35B MoE benchmarked locally on an RTX 3090 with 3B active params

gpjt · hn · 2026-07-25

A blog post benchmarks Qwen 3.6 35B MoE with 3B active parameters on an RTX 3090.

The focus is on how well a large MoE model can run locally on consumer hardware, making this more of a deployment and inference-cost piece than a pure model announcement. It is useful for people evaluating local setups and quantized performance.

Original post →

More from Infra

Infra channel →