Browser QA harness: GLM 5.3 Flash beats DeepSeek V4 Flash Vision on screenshots

Certain_Pension6305 · reddit · 2026-09-04

The author of a browser QA harness that turns screenshots into actions benchmarked DeepSeek V4 Flash Vision Exp vs GLM 5.3 Flash (both MIT-licensed). Access gap has narrowed: DeepSeek is now on Hugging Face via Novita or self-hosted via vLLM on one GB200 NVL4 tray; GLM has multiple hosted providers. Official scores aren't directly comparable (DeepSeek: 64.3 Chartography, 27.3 Agents Last Exam; GLM: 78.0 Chartography with Tools, 26.3 ALE) due to different harnesses. Running the same screenshot set through ZenMux with each model pinned, GLM gave clearly better overall results — its 30T-token multimodal pretraining likely explains the edge over DeepSeek's bolt-on vision modules. Verdict: GLM for this workload.

Original post →

More from Models

Models channel →