One DGX Spark, one ~180B MoE at 35 tok/s: 10k lines of code in 8 hours, fully local
Character-Result-281 · reddit · 2026-09-19
A developer built a project end-to-end on a single NVIDIA DGX Spark running Qwen3.8-Flash-Next (NVFP4, 262K context) fully locally:
- 8 hours of planning, coding and testing; 10k lines generated, 800k tokens consumed
- Stack: VSCode Copilot in autopilot mode + SGLang for inference
- The 180B MoE runs at 35 tok/s on the single box
The author admits it's not frontier-model quality, but demonstrates that local large-MoE full-stack development is now practical on workstation hardware.
More from coding & agent
- Anthropic's Head of Product Drops a 28-Minute Masterclass on Agents in Production — ifioknkem · 2026-09-20
- Teknium: Jev can't compact context well — Hermes summarizes 95% of it away — Teknium · 2026-09-20
- HarnessRouter open-sources a unified API to run Codex, Claude Code and more as agent backends — daniel_mac8 · 2026-09-20
- HarnessRouter: routing agent harnesses instead of models, a fresh infra idea — daniel_mac8 · 2026-09-20
- MCP tool naming: short generic verbs vs explicit prefixes for LLM tool selection — skvark · 2026-09-20
- GameToMac launched 10 days ago and already runs AoE IV, CS2 and Diablo IV on Apple Silicon — nickbaumann_ · 2026-09-20