Nemotron 3 Nano Omni Hits 264 tok/s Native on DGX Spark
ivan_bezdomny · x · 2026-08-03
A developer shared benchmarks running Nvidia's Nemotron-3-Nano-Omni-30B (3B active params) natively on a DGX Spark. The model achieves an impressive 264 tok/s inference and 1300 tok/s prefill at short contexts, dropping to 57 tok/s at 256k context.
Packed into a 33GB footprint, it integrates image, video, audio, OCR, and tool-calling capabilities. The author is currently wiring it into a ComfyUI workflow to build a fully local multimodal agent that can both see and generate images.
More from Infra
- A 10-Week Roadmap for LLM Inference Serving and Optimization — _jaydeepkarale · 2026-08-03
- App Developers Should Ship Their Own On-Device Models — abacaj · 2026-08-03
- Global AI compute to hit 200M H100-equivalents by 2028, fueling agentic loop toward ASI — 新智元 · 2026-08-03
- tinybox Dual-GPU Edition Hits 245 tok/s Running DeepSeek — AccBalanced · 2026-08-03
- Troubleshooting KV Cache Misses Caused by Multiple Agent Tool Calls — CentrifugalMalaise · 2026-08-03
- Wafer serves Kimi K3 on AMD MI355X with 3.8x throughput and 71% lower cost vs B200 — SumitGup · 2026-08-03