DeepSeek Models Tested: 242 One-shot Outputs Compared
kms_dev · reddit · 2026-08-03
Continuing a weekend of one-shotting cheap OpenRouter models, the author tested 10 DeepSeek models across 35 identical prompts.
Due to provider errors and empty completions, only 242 outputs made it out of the 350 matrix. The tested models include DeepSeek V4 (Pro/Flash), V3.x (V3.2/V3.1-terminus, etc.), and DeepSeek R1 series. An online comparison page is provided.
More from Models
- EpochAI Updates MirrorCode Leaderboard: Claude Fable 5 Leads with 64% Solve Rate — xeophon · 2026-08-04
- Are OpenAI/Anthropic Delaying Releases to Dodge Open Source Catch-up? — 0xsachi · 2026-08-04
- Running Frontier Models on 24GB VRAM: Local Deployment Challenges Cloud — mintybadgerme · 2026-08-04
- Rapid Thinking Beats Slow Intuition: RL Proves More Efficient Than Scaling Pretraining — intellectronica · 2026-08-04
- OpenAI Models Credited with 234 Math Results, Far Exceeding DeepMind — haider1 · 2026-08-04
- Unreleased OpenAI model solves 10 major math problems for $2,000 inference cost — Mobile_Distance_9598 · 2026-08-03