Leaked DeepSeek V4.1 benchmarks show 552B new-architecture model hitting 63.9 HLE with tools

teortaxesTex · x · 2026-09-10

A viral post shares what it claims are official DeepSeek V4.1 benchmarks (unconfirmed): a 552B total-parameter model with a new Causal-Encoder-Decoder architecture, 8B input / 16B output activation. Scores include 31.2 TerminalBench 4.0, 88.1 CyberGym, 15.3 ExploitGym, 63.9 HLE with tools, and 54.8 Automation-Bench — a major claimed leap in coding and agentic automation, pending official confirmation.

Related event: DeepSeek V4.1 and V4.1 Flash Benchmarks and Architecture Allegedly Leak(8 posts)→

Original post →

More from Models

Models channel →