Qwen3-VL plus DeCoDe lifts accuracy from 27.8% to 89.2% on novel vision tests

HildeKuehne · x · 2026-07-29

Qwen3-VL with DeCoDe jumps from 27.8% to 89.2% on a mixed benchmark suite

The post reports evaluation on 6 standard and 6 newly curated novel benchmarks, covering yoga poses, LEGO bricks, industrial parts, Egyptian hieroglyphs, insects, and Arabic sign language.

With Qwen3-VL + domain information, the DeCoDe method improves performance from 27.8% to 89.2%, and the authors say it outperforms SFT.

The result suggests that domain-aware decomposition can matter a lot when vision-language models face unfamiliar or highly specialized categories.

Original post →

More from Research

Research channel →