NVIDIA to bring Nemotron to RTX Spark, unveils DeepSeek V4 Flash running in just 60GB of memory
ryanshrout · x · 2026-10-08
NVIDIA's Ryan Shrout announced that a new Nemotron model will run on RTX Spark as part of a hybrid inference approach, and revealed a DeepSeek V4 Flash variant that needs only 60GB of memory. He also called for llama.cpp support on Windows ML.
More from Infra
- GitHub Copilot will offload tasks to local models like MAI-Code-1.1 Flash via HydraFusion — mariorod1 · 2026-10-08
- Microsoft launches Surface Spark with Nvidia: $6k for 128GB, odd form factor — casper_hansen_ · 2026-10-08
- Google launches first test satellite carrying 4 TPUs for its space-based ML infrastructure moonshot — CurieuxExplorer · 2026-10-08
- GPU rental platform Lium hits all-time-high utilization, courts idle GPU owners with 95%+ revenue share — const_reborn · 2026-10-08
- Marvell details Google chip deal worth up to $120B as inference accelerators split from TPUs — demian_ai · 2026-10-08
- llama.cpp Takes the Stage at Microsoft's Windows Event, Creator Celebrates Local AI Milestone — ggerganov · 2026-10-08