16-game Catan benchmark puts DeepSeek V4.1 Flash at GPT-5.6 Terra level at 6-7x lower cost

teortaxesTex · x · 2026-09-15

Developer @onusoz ran 16 games of Settlers of Catan pitting DeepSeek V4.1 Flash against GPT-5.6 Luna, Terra and Sol, with all runs open-sourced on Hugging Face (364 files, 741MB):

teortaxesTex called the comparison fair: in both subjective agentic coding and Catan, DeepSeek V4.1 Flash sits between Terra and Sol at a much lower price. He wants to see ARC-AGI 2 results and notes notable evaluators have been oddly indifferent to this model.

Related event: DeepSeek V4.1 Flash Matches GPT-5.6 in Catan, Aces Spatial Reasoning Test(2 posts)→

Original post →

More from Models

Models channel →