DeepSeek V4.1 Flash accused of public benchmark contamination, Kimi K3 possibly too
teortaxesTex · x · 2026-09-13
A user alleges DeepSeek V4.1 Flash shows benchmark contamination based on model-card data, with Kimi K3 possibly affected too. The claim: contaminated models rank disproportionately lower on later benchmarks unseen during training. Even if unintentional, the model's value reportedly lies in architectural innovation rather than raw output quality. Allegations unverified.
Related event: Debate Erupts Over Alleged Benchmark Contamination in DeepSeek V4.1 Flash(2 posts)→
More from Models
- GPT-6 Astra reportedly shipped 8 weeks after GPT-5.6 Sol, possibly the last fast turnaround — kimmonismus · 2026-09-13
- Dev fine-tunes Qwen 3.8 on 81,837 book annotations, writing quality up 86% in blind tests — Scobleizer · 2026-09-13
- Muse Spark 1.3 ties Fable 5.1 in 105-bug repo test, Meta joins frontier — PawelHuryn · 2026-09-13
- Frontier lab internal models reportedly lead public releases by 1-3 months — xeophon · 2026-09-13
- 45% of overnight benchmark rollouts failed mid-turn amid OpenAI capacity issues — dejavucoder · 2026-09-13
- Gemini turns out surprisingly good at translating swarm language to and from English — xeophon · 2026-09-13