DeepSeek Vision Test: Auto-fills forms but struggles with box alignment

teortaxesTex · x · 2026-08-22

A tester evaluated DeepSeek's newly released vision model using a form-filling benchmark. The model was required to locate and fill fields based solely on a screenshot, without access to coordinates or the DOM.

Methodology & Results:

Issues Identified:

Technical Insight: The reviewer notes the model uses a "foveated vision" approach (max 387 tokens per glance), designed to "look" rather than "see." The non-Exp version is expected to be more video-oriented, extracting sequences for reasoning.

Original post →

More from Multimodal

Multimodal channel →