Coverage tests across 3 runs: DeepSeek V4.1 Flash 21.7→33, GLM-5.3 Flash 17.7→24
PawelHuryn · x · 2026-09-15
Pawel Huryn shares more coverage test data across 3 runs: DeepSeek V4.1 Flash (max) improves from 21.7 to 33 points, while GLM-5.3 Flash (max) goes from 17.7 to 24 points. Third-party ongoing benchmark tracking of the two models.
More from Models
- OpenAI reportedly pauses $200 ChatGPT Pro signups as GPT-6 Astra demand saturates capacity — emmanuelvivier · 2026-09-15
- Vibe coding 3D game dev: Claude Pro beats ChatGPT Plus on limits and code audits — AdvertisingBubbly546 · 2026-09-15
- TerminalBench scoring isn't comparable: Astra gets zeros on safety stops, Fable re-routes to Opus — xeophon · 2026-09-15
- Quant Finance Shows What a Scaling-Pilled AI Industry Looks Like — and How the Moat Fades — willcb · 2026-09-15
- ChatGPT Plus vs Business: heavy users compare real-world Sol High usage limits — geler1 · 2026-09-15
- Tinfoil: Open-source safety classifiers are scarce, frontier labs urged to release more — yuntiandeng · 2026-09-15