Benchmark: pi agent harness outperforms DeepSeek's dsh on cost
solyarisoftware · x · 2026-08-20
DeepAPI benchmarked DeepSeek's official agent harness (dsh) against the pi framework.
- Test Scale: 180 controlled runs across 3 models and identical tasks.
- Cost: dsh was more expensive than pi on 2 out of 3 models.
- Analysis: The difference stemmed from round-trips; dsh required 11 model calls versus pi's 8.5 (on DeepSeek V4 Pro).
- Surprise: The only model where dsh won was Kimi K3, not DeepSeek.
Note: DeepSeek V4 Pro was run on GMICloud fp8, not the official API.
More from coding & agent
- Deep Agent + Stagehand: Browser Automation with Minimal Code — LangChain · 2026-08-20
- Claude Trading Skills open-sourced: Supports market analysis, risk management, and strategy development — tom_doerr · 2026-08-20
- Simon Willison: LOC metrics matter with agents, but conceptual integrity is harder to keep — Simon Willison · 2026-08-20
- Gemini Notebooks Integrates VM, Antigravity Coding Agent, and Skills Suite — AI_Andrew · 2026-08-20
- DeepSeekHarness RC.8: Multimodal input and ClaudeCode/Codex sub-agents — 机器之心 · 2026-08-20
- AI Agents Are Not Microservices: The Need for Durable Execution — rseroter · 2026-08-20