Blogger's test: DeepSeek V4.1 Flash scores 98% of GPT-6 Astra at 1.4% of the cost
solyarisoftware · x · 2026-09-10
Blogger NFTChen's everyday design-task benchmark puts DeepSeek V4.1 Flash at 81.2 vs GPT-6 Astra's 82.7 (98%), edging past Claude Fable 5.1 at 80.3 — at roughly 1.4% of Astra's cost ($0.023 vs $1.61) and about 2x faster (5.3 vs 11.1 minutes). A quoted cursive OCR test (270 chars) shows Kimi2.6 perfect, Qwen3.8-Max next, and DeepSeek V4.1 Flash tied for third (7 errors) with GLM 5.3 Flash and Gemini 3.1 Pro, but running at 260 tok/s with sub-3s single-image OCR and a recent price cut, making it strong for large-scale ancient-text digitization. Data from personal testing, not independently verified.
Related event: DeepSeek V4.1 Flash Matches 98% of GPT-6 Astra at 1.4% Cost(2 posts)→
More from Models
- Analyst: DeepSeek's latest change is a big win for token efficiency, moving toward OpenAI's regime — teortaxesTex · 2026-09-10
- antirez: stellar cybersec benchmarks aside, wait for real user testing before trusting the best scores — antirez · 2026-09-10
- antirez on the new DeepSeek model: 2-bit quants may not hold up, and it's not really 'local' — antirez · 2026-09-10
- DeepSeek report: post-training gains come from better data and environments, not RL novelty — realsohamparekh · 2026-09-10
- 'The whale is back': DeepSeek reportedly releases a new report — scaling01 · 2026-09-10
- Engram architecture explained: 500B backbone beats GLM 5.3-class with far smaller KV cache — bookwormengr · 2026-09-10