Frontier VLMs Still Fail at Parsing Forms; LlamaIndex Ships a Purpose-Built Fix

llama_index · x · 2026-09-29

LlamaIndex points out that even the latest frontier vision-language models (VLMs) still stumble on forms — because a form isn't plain text on a page, but a hierarchy of grouped fields, each tied to a specific box.

Their blog breaks down where VLMs fail on real W-2s, 1040s, W-9s and scanned W-4s, arguing forms need purpose-built parsing rather than a bigger general model: detect every field (not just obvious ones), preserve the section–field hierarchy, tie every value to its exact source box, and handle handwriting and checkmarks. An accompanying open-source cookbook shows how LlamaParse handles these at a fraction of the cost.

Original post →

More from coding & agent

coding & agent channel →