DeepSeek V4.1-Flash turns heads: GPT-5.6-level benchmarks at 552B params
teortaxesTex · x · 2026-09-10
DeepSeek released V4.1-Flash, with benchmarks reportedly at GPT-5.6 Sol level despite only 552B parameters, prompting "benchmarkmaxxed" skepticism. The architecture focuses on aggressive KV cache compression: a Causal Encoder-Decoder (CED) design projects decoder global KV from final encoder hidden states, activating just 8B params per token during prefill and 16B during decode—cost-effective for input-heavy agentic workloads. giffmana notes the native multimodal stack still uses SigLIP, suggesting a conservative choice.
More from Models
- Grok Voice Think Fast 2.0 High tops speech-to-speech leaderboard on task success — XFreeze · 2026-09-11
- Fable and Astra fail at the XY problem: eager executors with zero pushback and no metacognition — teodorio · 2026-09-11
- 'Chart Crime': Blogger Flags Cropped TerminalBench 4.0 Chart Underrating DeepSeek V4.1 — teortaxesTex · 2026-09-11
- Real API pricing method ranks subscriptions: GPT-5.6 Luna cheapest at $0.00083/MTok — teortaxesTex · 2026-09-11
- ThursdAI weekly: Astra, Navier Stokes, Muse assistant and more AI news — thursdai_pod · 2026-09-10
- DeepSeek v4.1-flash reportedly much better at subagent orchestration — adonis_singh · 2026-09-10