DeepSeek V4 Flash on dual local GPUs hits ~1400 tokens/sec fixing iOS bugs
HankYeomans · x · 2026-09-02
A developer ran DeepSeek V4 Flash (MXFP4, 1M context) locally via LM Studio as a coding agent and fixed two iOS bugs in one unattended run.
- Claude's harness caps 'unknown models' at 200K context, so he wired it up another way
- No CI, linting, or Xcode tests were set up; a few well-placed skills let it plan and solve the bugs, with Amp reviewing the work
- LM Studio logs show 1400 tokens/sec, consuming nearly all VRAM on two GPUs while still leaving 1.5GB for the Wayland display
- He calls the quality comparable to frontier coding agents in pre-testing
Related event: Dev Runs DeepSeek V4 Flash Locally on Dual GPUs to Fix iOS Bugs(2 posts)→
More from coding & agent
- Giving AI agents their own inbox is architecturally wrong, Reddit thread argues — Creamy-And-Crowded · 2026-09-02
- Merge launches enterprise AI governance tool enforcing model routing and spend rules — shensi · 2026-09-02
- Weaviate Ask Mode Adds Configurable Evaluation to Trade Latency vs Verifiability — CShorten30 · 2026-09-02
- Building the Ultimate Agent Harness for Kimi K3: The Model Is No Longer the Bottleneck — VibeMarketer_ · 2026-09-02
- GLM 5.2 slug references spotted in Google Antigravity CLI, hinting at integration — gaganghotra_ · 2026-09-02
- Design lead ships 12 PRs in a week: AI is erasing the designer-engineer gap — talkaboutdesign · 2026-09-02