Skipping a second RTX 5090 for two more Sparks: local inference user explains why

ideamaker321 · reddit · 2026-09-15

A local deployment enthusiast planned to buy a second RTX 5090 but opted for two more Spark units instead. The 5090 is fast for training small models, but running Qwen3.8 27B inference means constant optimizations and compaction issues—never a 'set and forget' experience. His existing two Sparks running Qwen Flash Next just work, and with two more he'll be able to run the larger DSF v4.1. He's asking others about their 5090 inference experiences.

Original post →

More from Infra

Infra channel →