DeepSeek 'King of Acing Benchmarks': Dev Says New Flash Model Falls Short of Hype
bindureddy · x · 2026-09-27
Developer Bindu Reddy calls DeepSeek the 'literal king of acing benchmarks,' mocking its latest Flash model's claim of matching Fable and Astra — arguing that a couple of minutes of real usage shows the claim doesn't hold up.
More from Models
- Internal benchmark: GLM-5.3 Flash writes scripts at 1/429th the price, 1.7 points behind — OnlyProggingForFun · 2026-09-27
- Grok accused of uploading user chat images to the web as Musk says 'this keeps getting worse' — EthanJPerez · 2026-09-27
- Open-source 'Jev-like' models proliferate before anyone agrees on a definition — MaziyarPanahi · 2026-09-27
- Xiaomi's MIT-licensed 1.02T MoE MiMo-V2.6-Pro ties Grok 4.7 at 46 on Artificial Analysis — dl_weekly · 2026-09-27
- Dev Impressed: 'Opus 5.5 Is Good at Games' in Early Hands-On Claim — tobowers · 2026-09-27
- AMA: 167 models charted on a creative-writing benchmark Pareto frontier — OnlyProggingForFun · 2026-09-27