DeepSeek V4.1 Flash scores 40 in AA index despite rising hallucinations
Artificial Analysis gives DeepSeek V4.1 Flash a score of 40, beating DeepSeek V4 Pro though trailing GLM-5.3-Flash. Its AA-Omniscience rose 9 points to -5.3, but hallucination rates increased.
2026-09-11 ~ 2026-09-11 · 2 related posts
- Episode 1: DeepSeek V4.1 Flash Opens Limited Internal Beta with New Architecture and Native Multimodality(2026-09-08, 18 posts)
- Episode 2: DeepSeek V4.1 Flash Tested: Blazing 350 Tokens/s but Still Experimental(2026-09-08, 2 posts)
- Episode 3: DeepSeek cuts V4-Flash API prices with new peak/off-peak billing from Sept 10(2026-09-08, 6 posts)
- Episode 4: DeepSeek V4.1 Flash tested across tasks: near-frontier performance at a fraction of the cost(2026-09-09, 15 posts)
- Episode 5: DeepSeek releases V4.1-Flash: 552B MoE beats flagships at low cost(2026-09-09, 56 posts)
- Episode 6: DeepSeek V4.1 Flash Leak: 552B Asymmetric MoE Reportedly Rivals GPT-5.6(2026-09-10, 20 posts)
- Episode 7: DeepSeek V4.1 Flash Benchmark Results Leak, Sparking Community Discussion(2026-09-10, 2 posts)
- Episode 8: Bug Hunt Bench: DeepSeek V4.1 Flash Tops Price-Performance(2026-09-10, 8 posts)
- Episode 9: Inside DeepSeek V4.1 Flash: YOCO at its core and KV cache reuse(2026-09-10, 7 posts)
- Episode 10: DeepSeek V4.1 Flash Deep Dive: KV Cache Compression at the Frontier(2026-09-11, 9 posts)
- Episode 11: DeepSeek's New Model Report: 4x Smaller KV Cache, More Stable Training(2026-09-11, 3 posts)
- Episode 12: DeepSeek V4.1 Flash tops Vals open-source index at $0.30 per run(2026-09-11, 2 posts)
- Episode 13: DeepSeek V4.1 Flash scores 40 in AA index despite rising hallucinations(2026-09-11, 2 posts)
- Episode 14: Inside KV-Cache Sharing: How DeepSeek CED and GLM 5.2 Differ(2026-09-11, 4 posts)
- DeepSeek-V4.1-Flash scores 40 on Artificial Analysis Index, beating DeepSeek-V4-Pro — scaling01 · 2026-09-11
- DeepSeek V4.1 Flash evals: accuracy up but non-hallucination rate falls — ArtificialAnlys · 2026-09-11