Dev Runs DeepSeek V4 Flash Locally on Dual GPUs to Fix iOS Bugs

A developer reported running DeepSeek V4 Flash locally with MXFP4 quantization and 1M context on dual GPUs, using it as a coding agent to plan and fix two iOS bugs at around 1400 tokens per second.

2026-09-02 ~ 2026-09-02 · 2 related posts

1 near-duplicate retellings: HankYeomans