DeepSeek v4 Pro on 4×GB300: Sub-Agent Hits 188 tok/s
Xianbao_QIAN · x · 2026-07-03
A user deployed DeepSeek v4 Pro + DSpark on 4 GB300 GPUs, integrating it into opencode with default settings. Tests showed each sub-agent achieving a blistering 188 tok/s. The author described the coding experience as returning to the intuitive feel of native JS. Expressing immense excitement over dedicated inference compute, the author is even considering purchasing GB300 hardware to fully embrace an unlimited, dedicated agent lifestyle without waiting for slow LLM responses.
More from coding & agent
- Tweaked orchestration skill turns agents into self-policing workflow — pvncher · 2026-07-27
- A practical map of 11 protocols in the modern AI agent stack — TheTuringPost · 2026-07-27
- Qwen Code nightly adds Goal v3 orchestration and workspace channel controls — qwen-code-ci-bot · 2026-07-27
- NVIDIA says Nemotron 3 Ultra hit 97.1% on agentic RTL chip-design tasks — NVIDIAAI · 2026-07-27
- Tokyo Agent Forge hackathon shipped production-ready AI agents in one day — DavidBennett__ · 2026-07-27
- Long-running agents will need immutable event logs, this thread argues — sebpaquet · 2026-07-27