Browser QA harness: GLM 5.3 Flash beats DeepSeek V4 Flash Vision on screenshots
Certain_Pension6305 · reddit · 2026-09-04
The author of a browser QA harness that turns screenshots into actions benchmarked DeepSeek V4 Flash Vision Exp vs GLM 5.3 Flash (both MIT-licensed). Access gap has narrowed: DeepSeek is now on Hugging Face via Novita or self-hosted via vLLM on one GB200 NVL4 tray; GLM has multiple hosted providers. Official scores aren't directly comparable (DeepSeek: 64.3 Chartography, 27.3 Agents Last Exam; GLM: 78.0 Chartography with Tools, 26.3 ALE) due to different harnesses. Running the same screenshot set through ZenMux with each model pinned, GLM gave clearly better overall results — its 30T-token multimodal pretraining likely explains the edge over DeepSeek's bolt-on vision modules. Verdict: GLM for this workload.
More from Models
- Gemini 3.8 Flash bug fixed: Google AI Mode now shows far more source links — gaganghotra_ · 2026-09-04
- Critics warn OpenAI's GPT-6 Astra reasons opaquely, gutting CoT monitoring safety — GaryMarcus · 2026-09-04
- ChatGPT has a 'cyber abuse' ban reason: pushing the model too hard gets you banned — sven_ai · 2026-09-04
- ChatGPT adds writing-style matching from connected apps, analytics, and a Yubikey deal tied to Daybreak access — btibor91 · 2026-09-04
- GPT-6 reportedly launches as Tesla starts public rides in steering-free Cybercab — Dr_Singularity · 2026-09-04
- Astra early-access users' similar blender demos look coordinated, with no practical examples shown — jdjohnson · 2026-09-04