Leaked DeepSeek V4.1 benchmarks show 552B new-architecture model hitting 63.9 HLE with tools
teortaxesTex · x · 2026-09-10
A viral post shares what it claims are official DeepSeek V4.1 benchmarks (unconfirmed): a 552B total-parameter model with a new Causal-Encoder-Decoder architecture, 8B input / 16B output activation. Scores include 31.2 TerminalBench 4.0, 88.1 CyberGym, 15.3 ExploitGym, 63.9 HLE with tools, and 54.8 Automation-Bench — a major claimed leap in coding and agentic automation, pending official confirmation.
Related event: DeepSeek V4.1 and V4.1 Flash Benchmarks and Architecture Allegedly Leak(8 posts)→
More from Models
- vLLM Ships Full Support for DeepSeek-V4.1-Flash's New Architecture — vllm_project · 2026-09-10
- DeepSeek-V4.1-Flash reportedly offers continuous reasoning effort from 1 to 100 — zainhas · 2026-09-10
- DeepSeek multimodal team praised for unusually careful pre-training data work — zephyr_z9 · 2026-09-10
- DeepSeek v4.1 flash evals leak: SOTA on deepswe, competitive on terminal bench — zainhas · 2026-09-10
- A new kind of encoder-decoder model has appeared, researcher says — bclavie · 2026-09-10
- DeepSeek V4.1 Flash has just 20 decoder layers, shallower than its predecessor — teortaxesTex · 2026-09-10