Document classification bake-off: VLM is ~7x more token-efficient than OCR
spillai · x · 2026-10-02
VLM Run ran a bake-off comparing VLM-based document classification against OCR-based classification using their TypeSafe-compatible SystemOne API:
- Token efficiency: the VLM path consumes a fixed 372 tokens per page, while OCR->Text medians around 2,002 tokens with a worst case of 19,159 on a single page — over 7x more efficient for VLM.
- The core finding: VLMs can be surprisingly accurate and token-efficient on the right tasks, and their fixed consumption makes capacity planning predictable.
First-party benchmark data, directly useful for teams building document processing pipelines.
More from coding & agent
- steipete: btrfs CoW is great for worktrees but terrible for sqlite — steipete · 2026-10-02
- Google Mantis: a skills pack for security review with coding agents — udmrzn · 2026-10-02
- NVIDIA and Cedana serve a portfolio of coding models on one 8x B200 node — josh_wills · 2026-10-02
- LukeW teases Intent Mobile: agent coordination at the level of intent — LukeW · 2026-10-02
- LukeW: Dev tools haven't reached their final form after Terminal UI and chat clients — LukeW · 2026-10-02
- Dev Benchmarks 10+ Coding Agent Workflows: Astra Planning + Sol 6.1 Beats Pricier Combos — kevinkern · 2026-10-02