Qwen3-VL 8B on a MacBook Beats GPT-5.6 on Tax Forms, 89% Opus Tops 137-Doc Benchmark
NegotiationKey7184 · reddit · 2026-09-28
The author benchmarked Qwen3-VL 8B Instruct (Q4KM via Ollama, M5 MacBook 24GB, 30s/doc) against Claude Opus 5.5, Sonnet 5 and GPT-5.6 Terra on 137 messy documents: CORD/SROIE receipts, 20 scanned 1980s-90s invoices, 32 freshly generated IRS forms, 10 Indian bank statements, and 15 CUAD contracts.
Fully-correct scores: Opus 89%, Sonnet 85%, Qwen 8B 59%, GPT-5.6 Terra 57%. Notable findings:
- W-2s: Qwen 21/32 fully correct vs GPT-5.6 Terra's 7/32
- Indian statements: 2/10—all amounts right but dd-mm-yyyy dates read as mm-dd
- Long contracts: 2/15, mostly wrong expiry dates
- GPT-5.6 Terra silently "corrects" unusual spellings (Rachael→Rachel)
- Self-checking barely helps: 119/137 outputs unchanged
- At least 4 published SROIE answer keys are wrong
- Ollama's default qwen3-vl:8b is the thinking variant that ignores think:false; use :8b-instruct
The author plans to fine-tune the 8B to fix date/spelling failures. Prompts, keys, scorers and raw outputs are open-sourced on GitHub (messy-docs-bench).
More from Models
- DeepMind's Veo/Gemini Omni lead rebuilds his site with Antigravity—'it just works' — dumierhan · 2026-09-28
- OpenAI Pauses Training Its Most Powerful Models After Rogue Agents Target Government — wiredmagazine · 2026-09-28
- VeriLoop E2 3-bit quant beats Qwen 27B on real website-generation task — zyxciss · 2026-09-28
- User: Gemini integration in Google Maps never responds to questions — BlackHC · 2026-09-28
- H Company releases Holo4 open VLMs for computer-use agents — jacek2023 · 2026-09-28
- User ditches Gemini Astra for Claude Opus: 5x more work done, results nail prompts on first try — Historical_Buyer5248 · 2026-09-28