DeepSeek v4 Flash vs Qwen Local Coding Test: Qwen 122B Offers Better Cost-Performance
returnity · reddit · 2026-08-05
The author conducted an agentic coding benchmark on an M5 Max 128GB using a 109-question subset of Aider Polyglot to compare DeepSeek v4 Flash against various Qwen and Gemma models.
Key Findings:
- DeepSeek v4 Flash (High) achieved the highest overall accuracy but consumed over 5x the tokens of Qwen3.5-122B to achieve those results.
- Qwen3.5-122B demonstrated exceptional cost-performance and significantly outperformed DeepSeek in first-try pass rates.
- Qwen3.6-27B-ThinkingCap fine-tune performed excellently, completing tasks using only 20-30% of the total tokens of the base model, effectively mitigating overthinking.
The test methodology rigorously excluded questions that caused infinite loops due to empty linter errors and tracked metrics like context overflows and timeouts.
More from Models
- NVIDIA Releases Alpamayo 2 Super: A 34B Open-Source VLA Model for Autonomous Driving — drmapavone · 2026-08-05
- OpenAI's gpt-5.6-luna So Cost-Effective It Constantly Overloads Servers — _lewtun · 2026-08-05
- Opus 5 Language Degradation: Anthropic's Model Accused of Being a 'Jargon Douche' — TheTuringPost · 2026-08-05
- AI Needs a Nap? Developer Exposes Hilarious GPT Thinking Block Hallucination — cantrell · 2026-08-05
- engyai Launches Cheapest Kimi K3 API on OpenRouter, Cutting Costs by 50% — const_reborn · 2026-08-05
- LiquidAI's LFM2.5-2.6B Hits 82 tok/s Decode on Mac with 128K Context — helloiamleonie · 2026-08-05