Qwen3-VL 8B on a MacBook Beats GPT-5.6 on Tax Forms, 89% Opus Tops 137-Doc Benchmark

NegotiationKey7184 · reddit · 2026-09-28

The author benchmarked Qwen3-VL 8B Instruct (Q4KM via Ollama, M5 MacBook 24GB, 30s/doc) against Claude Opus 5.5, Sonnet 5 and GPT-5.6 Terra on 137 messy documents: CORD/SROIE receipts, 20 scanned 1980s-90s invoices, 32 freshly generated IRS forms, 10 Indian bank statements, and 15 CUAD contracts.

Fully-correct scores: Opus 89%, Sonnet 85%, Qwen 8B 59%, GPT-5.6 Terra 57%. Notable findings:

The author plans to fine-tune the 8B to fix date/spelling failures. Prompts, keys, scorers and raw outputs are open-sourced on GitHub (messy-docs-bench).

Original post →

More from Models

Models channel →