The best local model you can run on 2 GB10s, per this desk setup
jasonkneen · x · 2026-09-06
A user showcases the best local model setup currently runnable on two GB10 boxes sitting on his desk, sharing the actual hardware configuration. Useful reference for local deployment enthusiasts.
More from Infra
- A First-Principles Handbook on KV Cache: From MHA/GQA/MLA to PagedAttention — techNmak · 2026-09-06
- One ComfyUI node fixed MiniMax H3 OOM on RTX 5090: full workflow for 15s 2K video — denizbuyukayak · 2026-09-06
- KV cache often spills out of HBM in the agentic era, tanking effective bandwidth — AccBalanced · 2026-09-06
- Hybrid bonded HBM hypothetical market: over 3 billion D2D applications per year — zephyr_z9 · 2026-09-06
- Ollama CEO: open models will carry 80-90% of enterprise tokens at just 10-20% of cost — victor_explore · 2026-09-06
- Nvidia de-specced Rubin Ultra HBM from 12-Hi to 8-Hi: $/bandwidth is the bottleneck — AccBalanced · 2026-09-06