Real-world agent eval: DeepSeek V4.1 Flash beats Haiku 5.5 on both tasks

dergachoff · reddit · 2026-10-08

The author benchmarked Haiku 5.5 against DeepSeek V4.1 Flash on two real agent jobs in their app; DeepSeek won both.

Job 1 — research sub-agent (4 real briefs, Exa/Brave search + fetcher + image analysis):

Job 2 — chat titles (52 messages, reasoning off): DeepSeek won 28 vs Haiku 8. Same speed (1.1s), Haiku slightly cheaper. Haiku often answers the message instead of titling it, and one prompt-injection test message became a title.

Method: Opus judged blind A/B with reversed orders; ties when orders disagreed. Acknowledged same-lab bias (Claude judging Claude), yet it still picked DeepSeek. Single-run small sample, only indicative for the author's own workloads.

Related event: Hands-on Tests Show DeepSeek V4.1 Flash Beats Haiku 5.5 on Agent Tasks(2 posts)→

Original post →

More from coding & agent

coding & agent channel →