Running Agentic Coding on Two DGX Sparks: 100M Tokens/Day Economics

HarveenChadha · x · 2026-08-02

A developer shared empirical estimates for running agentic coding workloads on local hardware. Using two DGX Sparks to run original model weights, it can achieve 80-90 tokens/s in a single stream at an 85% cache hit rate.

Under these conditions, this setup can process up to 100 million input tokens and generate 6 million output tokens per day. The author asks the community if this $9,500 hardware configuration is the right setup.

Related event: DeepSeek-V4-Flash Local Deployment Benchmarks: Performance Across Hardware(21 posts)→

Original post →

More from coding & agent

coding & agent channel →