Skipping a second RTX 5090 for two more Sparks: local inference user explains why
ideamaker321 · reddit · 2026-09-15
A local deployment enthusiast planned to buy a second RTX 5090 but opted for two more Spark units instead. The 5090 is fast for training small models, but running Qwen3.8 27B inference means constant optimizations and compaction issues—never a 'set and forget' experience. His existing two Sparks running Qwen Flash Next just work, and with two more he'll be able to run the larger DSF v4.1. He's asking others about their 5090 inference experiences.
More from Infra
- Liam Fedus' Lab Built Neon: 1,300 H200s and a Self-Driving Materials Loop Beat GPT-6 Astra at Superconductor Analysis — vwxyzjn · 2026-09-16
- Periodic Labs: 1,300 H200s beat frontier models on X-ray evals, 4.1x Megatron throughput — zephyr_z9 · 2026-09-16
- NVIDIA and Palantir plant a flag for sovereign AI, starting with NVIDIA — postalex · 2026-09-16
- Ben Bajarin bullish on Credo: DustPhotonics deal and vertical integration undervalued by market — BenBajarin · 2026-09-16
- Unconventional AI shows analog oscillator chip, claims 1000x power efficiency over Nvidia in 2 years — PTrubey · 2026-09-16
- RAM-Pooling Across 3 Old Devices Runs a 40B Model at 16 tok/s — With a Counterintuitive CPU Finding — Medicine_Blogscanner · 2026-09-16