Nemotron 3 Nano Omni Hits 264 tok/s Native on DGX Spark

ivan_bezdomny · x · 2026-08-03

A developer shared benchmarks running Nvidia's Nemotron-3-Nano-Omni-30B (3B active params) natively on a DGX Spark. The model achieves an impressive 264 tok/s inference and 1300 tok/s prefill at short contexts, dropping to 57 tok/s at 256k context.

Packed into a 33GB footprint, it integrates image, video, audio, OCR, and tool-calling capabilities. The author is currently wiring it into a ComfyUI workflow to build a fully local multimodal agent that can both see and generate images.

Original post →

More from Infra

Infra channel →