Real-world agent eval: DeepSeek V4.1 Flash beats Haiku 5.5 on both tasks
dergachoff · reddit · 2026-10-08
The author benchmarked Haiku 5.5 against DeepSeek V4.1 Flash on two real agent jobs in their app; DeepSeek won both.
Job 1 — research sub-agent (4 real briefs, Exa/Brave search + fetcher + image analysis):
- Haiku low: 1W/3L, 480s, $0.09; medium: 1/1/2; high: 0W/3L, 992s, $0.31
- DeepSeek low: 3W/1L, 721s, $0.35
- Haiku is faster and 4x cheaper but consistently lost two briefs: DeepSeek found the brand's own guideline page with exact HEX/Pantone colors, while Haiku never did even at high effort with 60 tool calls — higher effort meant more searching, not better finding. Its only win (forum quotes) relied solely on Exa snippets without opening pages.
Job 2 — chat titles (52 messages, reasoning off): DeepSeek won 28 vs Haiku 8. Same speed (1.1s), Haiku slightly cheaper. Haiku often answers the message instead of titling it, and one prompt-injection test message became a title.
Method: Opus judged blind A/B with reversed orders; ties when orders disagreed. Acknowledged same-lab bias (Claude judging Claude), yet it still picked DeepSeek. Single-run small sample, only indicative for the author's own workloads.
Related event: Hands-on Tests Show DeepSeek V4.1 Flash Beats Haiku 5.5 on Agent Tasks(2 posts)→
More from coding & agent
- Dev compares vibe coding to buying books you never read: 15-min PoCs, zero learning — MarcJSchmidt · 2026-10-08
- Blockchain instructor of 5 years: don't learn to code, learn vibe coding and product design — mfckr_eth · 2026-10-08
- Solo dev shares the AI stack behind a $1M MRR portfolio for less than a junior hire's week — tibo_maker · 2026-10-08
- React Doctor by Aiden Bai wins rave reviews from agent-powered developers — aidenybai · 2026-10-08
- Running 4 Claude agents with zero guardrails: a power user questions local agent sandboxing — Adventurous-Simple99 · 2026-10-08
- A production-grade repo blueprint for LLM apps, from API layers to evals — techNmak · 2026-10-08