Dev Runs DeepSeek V4 Flash Locally at ~1400 tok/s, Fixes Two iOS Bugs End-to-End
HankYeomans · x · 2026-09-02
Developer HankYeomans tested a locally deployed DeepSeek V4 Flash (MXFP4, claimed 1M context) as a coding agent.
- Found the Claude harness caps 'unknown models' at 200K context, so he ran it via LM Studio instead.
- After joint planning, the model autonomously executed phases and handled commits end-to-end; with a few skills it fixed two iOS bugs without CI, linting, or Xcode testing set up.
- It nearly maxed out two GPUs' VRAM; LM Studio logs showed 1400 tokens/sec, and he had Amp (Ultra) review the work.
- He says the experience rivals pre-release frontier coding agents. Note: DeepSeek V4 Flash is not yet officially confirmed.
Related event: Dev Runs DeepSeek V4 Flash Locally on Dual GPUs to Fix iOS Bugs(2 posts)→
More from coding & agent
- Anthropic ships official Claude Fable 5.1 prompting guide with 16 fixes for agent pain points — xiaohu · 2026-09-02
- Claude Fable 5.1 prompting guide: 16 official tips and an agent migration checklist — xiaohu · 2026-09-02
- Polar Analytics launches Polar Operator, an AI operator for commerce teams in Slack — JaynitMakwana · 2026-09-02
- $40K of guitar gear won't make you Mike Rutherford — why coding skill still wins with AI agents — JFPuget · 2026-09-02
- Agents-run SEO pipeline: 8 draft PRs in two weeks, most runs under $1 — Pitiful-Surround-285 · 2026-09-02
- Vibe coding beats product hunting: build the exact tool you need — Philmod · 2026-09-02