DeepSeek V4.1 Flash beats Haiku 5.5 in dev's small eval on research sub-agent and title tasks
dergachoff · reddit · 2026-10-08
A GPU-poor developer (M2 Max, 32GB) tested whether Haiku 5.5 could replace DeepSeek V4.1 Flash in his app. It couldn't.
Job 1, research sub-agent (web search + page reading + sourced reports, same Exa/Brave tools): Haiku low went 1W/3L vs DeepSeek (480s, $0.09 vs 721s, $0.35), medium 1W/1T/2L, high 0W/1T/3L. Haiku is up to 4x cheaper but lost the same two briefs every time — notably one requiring the brand's own guideline page with exact HEX/Pantone colors, which DeepSeek found and Haiku never did even at high effort with 60 tool calls. Haiku's only win used Exa snippets without opening pages, of dubious provenance.
Job 2, chat titles (52 messages, reasoning off): DeepSeek 28 wins, Haiku 8, 9 splits, 7 identical, same 1.1s speed. Haiku often answered the message instead of titling it, and one prompt-injection test message literally became the title.
Method: blind A/B judging by Opus with reversed orders; author acknowledges same-lab bias, yet it still picked DeepSeek. Single runs per setup — a personal eval, not a lab benchmark.
Related event: Hands-on Tests Show DeepSeek V4.1 Flash Beats Haiku 5.5 on Agent Tasks(2 posts)→
More from Models
- Unsloth trains local decision models on 3GB VRAM, lifting accuracy from 30% to 78% — evilsocket · 2026-10-08
- OpenAI researcher: Claude coulda solved the Quasi-Riemann Hypothesis — burny_tech · 2026-10-08
- High-signal critique: GPT 5.4 lacks original moves, autonomous superhacker 'is not there' — teortaxesTex · 2026-10-08
- ThePrimeagen: It's worse than I thought — OpenAI one-shotted a math problem — burny_tech · 2026-10-08
- Bug Hunt Bench: blind-graded leaderboard testing frontier coding models on 105 real planted bugs — PawelHuryn · 2026-10-08
- Sonnet 5.5 wins by running 7x more turns than GPT-6 Astra, but costs most and is slowest — PawelHuryn · 2026-10-08