DeepSeek Vision Test: Auto-fills forms but struggles with box alignment
teortaxesTex · x · 2026-08-22
A tester evaluated DeepSeek's newly released vision model using a form-filling benchmark. The model was required to locate and fill fields based solely on a screenshot, without access to coordinates or the DOM.
Methodology & Results:
- Task: Identify form fields and simulate clicks/inputs using only visual data.
- Performance: The model probed positions, took screenshots of itself, placed elements, and nudged them into place. It filled every field and performed self-verification before submission.
- Stats: Took 5 minutes 48 seconds; consumed 288k input tokens and 38k output tokens.
Issues Identified:
- Checkmark ticks landed next to the box rather than inside it.
- Character-box drift occurred with longer text strings.
Technical Insight: The reviewer notes the model uses a "foveated vision" approach (max 387 tokens per glance), designed to "look" rather than "see." The non-Exp version is expected to be more video-oriented, extracting sequences for reasoning.
More from Multimodal
- ZastTranslate: Local video translation and dubbing with voice cloning — cocktailpeanut · 2026-08-22
- Midjourney style code shared: sref 2564258308 — sergeantsref · 2026-08-22
- Grok 4.6 cached input tokens offer only 75% discount, below standard 90% — cocktailpeanut · 2026-08-22
- AI demo showcases generated animated CAD assemblies — jakedahn · 2026-08-22
- Opinion: AI Shifts Music Creation from Training to Taste — thisguyknowsai · 2026-08-22
- Fully AI-Generated Police Procedural 'Dick Snifford' Debuts — venturetwins · 2026-08-22