Qwen3-VL plus DeCoDe lifts accuracy from 27.8% to 89.2% on novel vision tests
HildeKuehne · x · 2026-07-29
Qwen3-VL with DeCoDe jumps from 27.8% to 89.2% on a mixed benchmark suite
The post reports evaluation on 6 standard and 6 newly curated novel benchmarks, covering yoga poses, LEGO bricks, industrial parts, Egyptian hieroglyphs, insects, and Arabic sign language.
With Qwen3-VL + domain information, the DeCoDe method improves performance from 27.8% to 89.2%, and the authors say it outperforms SFT.
The result suggests that domain-aware decomposition can matter a lot when vision-language models face unfamiliar or highly specialized categories.
More from Research
- LLMs on robots jump real-world success from 16.7% to 97.3% — tri_dao · 2026-07-29
- Simple teleoperation recordings plus soft compliance can solve more robot tasks than expected — ihorbeaver · 2026-07-29
- Studies find AI therapy replies often score higher on empathy than human clinicians — sapinker · 2026-07-29
- TransluceAI proposes oversight foundation models to catch reward hacking at scale — JacobSteinhardt · 2026-07-29
- Student builds a local coding agent on molab and beats two 7B coder baselines — S_Conradi · 2026-07-29
- Why tokenizer optimization is hard: expensive pretraining, slow feedback, and non-differentiable design — paul_cal · 2026-07-29