'Chart Crime': Blogger Flags Cropped TerminalBench 4.0 Chart Underrating DeepSeek V4.1
teortaxesTex · x · 2026-09-11
teortaxesTex mocks a TerminalBench 4.0 chart as "chart crime": the blog post's figure shows TerminalBench 4.0 and DeepSWE scoring lower than DeepSeek V4.1. The quoted post points out the blog screenshot was cropped and links the full original image. A dispute over how DeepSeek's new model is presented on benchmarks.
More from Models
- GPT-6 "Sol" reference spotted online, release rumored ahead of DevDay — koltregaskes · 2026-09-11
- Ex-OpenAI staffer claims frontier models are finetuned on your successful chats — burkov · 2026-09-11
- Teknium claims DeepSeek Flash V4.1 is live in Hermes Agent, unverified — Teknium · 2026-09-11
- OpenAI's Math Breakthroughs Now Face Default Skepticism Over 'Stolen Work' Claims — SGC-UNIT-555 · 2026-09-11
- Rumors Swirl: OpenAI Near Verifying Hodge Conjecture Proof, BSD May Be Next — TheGoldenLeaper · 2026-09-11
- Goodfire finds a general-purpose addition module in Llama 3.1 8B — burny_tech · 2026-09-11