DeepSeek accused of benchmark gaming: new Flash model underwhelms in real use
Arindam_1729 · x · 2026-09-27
@bindureddy called DeepSeek "the king of acing benchmarks," arguing its new Flash model claims parity with Fable and Astra but falls apart after two minutes of actual use. Replying, Arindam noted the model's vision capabilities do hold up well in his testing.
Related event: Developers Slam DeepSeek's New Model as Benchmark Champ, Real-World Flop(2 posts)→
More from Models
- Martin Casado recommends the best talk on in-context learning, a first-principles view of LLMs — AccBalanced · 2026-09-27
- Hands-on: Opus 5.5 high beats GPT-6 astra xhigh on real Pagespeed optimization — mazzaTalk · 2026-09-27
- Dev on Opus 5.5: smooth multi-part coordination flips coding dynamic — ezshine · 2026-09-27
- Asked Grok to teach Japanese kanji, it started inventing its own — JoeJustice · 2026-09-27
- GPT Agents Are Writing Eerie Uncategorizable Short Stories, and This One Is a Gem — RileyRalmuto · 2026-09-27
- Google ships a UI in AI Studio for trying its new audio and voice models — ammaar · 2026-09-27