How to Top the Decision Index Vision: Remove the 512x512 Cap in the HF Implementation

antoine_chaffin · x · 2026-10-09

The author shares how they reached #1 on the Decision Index Vision: examine the data, notice underperformance on tasks requiring precise in-image reading, then connect the dots — the HF implementation caps inputs at 512x512 resolution. Raising the cap delivered the win.

PSA: if you're using their decider for vision tasks via the HF model, raise the max resolution yourself. Thanks to the Qwen backbone, the model is likely already near SOTA on vision tasks. The API is unaffected and already supports 2M-pixel images.

Related event: Decision Index Vision Winner Found HF's 512-Resolution Cap Was the Bottleneck(2 posts)→

Original post →

More from Models

Models channel →