NVIDIA to bring Nemotron to RTX Spark, unveils DeepSeek V4 Flash running in just 60GB of memory

ryanshrout · x · 2026-10-08

NVIDIA's Ryan Shrout announced that a new Nemotron model will run on RTX Spark as part of a hybrid inference approach, and revealed a DeepSeek V4 Flash variant that needs only 60GB of memory. He also called for llama.cpp support on Windows ML.

Related event: Microsoft and Nvidia unveil RTX Spark as Windows pivots to local AI at Surface event(14 posts)→

Original post →

More from Infra

Infra channel →