DeepSeek v4.1 flash evals leak: SOTA on deepswe, competitive on terminal bench
zainhas · x · 2026-09-10
An unverified look at DeepSeek's new v4.1 flash eval numbers: reportedly SOTA on deepswe alongside sol and fable5, and very competitive on terminal bench 3. Official details are still pending.
Related event: DeepSeek V4.1 Flash Benchmark Results Leak, Sparking Community Discussion(2 posts)→
More from Models
- Leaked brief: DeepSeek V4.1 Flash at 552B params, foldable iPhone at $1,999, ChatGPT voice limits raised — testingcatalog · 2026-09-10
- CoT similarity test suggests Qwen3.8 may have been trained on GPT 5.5 reasoning traces — Chromix_ · 2026-09-10
- DeepSeek accused of 'pretending linear attention doesn't exist' in new architecture — teortaxesTex · 2026-09-10
- Hands-on with GLM 5.3 Flash: great at coding and research, poor at trading and ideation — ManagementNo5153 · 2026-09-10
- Tell an agent it has 1M token budget and it reasons 3-5x longer: a metacognition experiment — paraschopra · 2026-09-10
- DeepSeek's new open model beats GLM 5.3 and Kimi K3 at 4-10x lower price — deedydas · 2026-09-10