16-game Catan benchmark puts DeepSeek V4.1 Flash at GPT-5.6 Terra level at 6-7x lower cost
teortaxesTex · x · 2026-09-15
Developer @onusoz ran 16 games of Settlers of Catan pitting DeepSeek V4.1 Flash against GPT-5.6 Luna, Terra and Sol, with all runs open-sourced on Hugging Face (364 files, 741MB):
- DeepSeek easily beat Luna; finished games against Terra ended in a 4-4 tie while costing 6-7x less
- Sol generally won, but DeepSeek still took one game off it; Luna had no chance against DeepSeek
- Games with both models on max thinking took too long to finalize
- Models ran via Novita through Hugging Face Inference Providers
teortaxesTex called the comparison fair: in both subjective agentic coding and Catan, DeepSeek V4.1 Flash sits between Terra and Sol at a much lower price. He wants to see ARC-AGI 2 results and notes notable evaluators have been oddly indifferent to this model.
Related event: DeepSeek V4.1 Flash Matches GPT-5.6 in Catan, Aces Spatial Reasoning Test(2 posts)→
More from Models
- AI Overview auto-applies the 'creative writing' jailbreak, researcher shows — conitzer · 2026-09-15
- Solar Pro 4 stays free on Nous Portal until Sept 25 in Upstage partnership — keunwoochoi · 2026-09-15
- Bug Hunt Bench: multiple runs boost small-model bug detection but move frontier models just 1-2 points — PawelHuryn · 2026-09-15
- ZDTaichu5.0-9B, a 9B vision-language model with spatial reasoning, trends on Hugging Face — TaichuAI · 2026-09-15
- Which 10Eros video-model quant works best on 8GB VRAM? A practical trade-off question — apostrophefee · 2026-09-15
- rasbt shows why final-result benchmarks mislead: Astra vs Qwen in Paint — rasbt · 2026-09-15