214 API calls later: Ling-3.0-flash-VL spots a 30-yuan receipt error and shines on video detail
alifcoder · x · 2026-09-23
A systematic hands-on test of Ling-3.0-flash-VL across 214 API calls covers image and video understanding. In the standout case, the author fed it a 900×870 receipt with 12 line items and a printed total of 2,092.10, using a three-step prompt: compute the real total, compare with the printed value, then judge consistency — the model itemized all lines in 6 seconds and caught a 30-yuan discrepancy.
Short-video detail recognition also performed well. The tests probed edge cases including small-text OCR, counting, and chart reading.
More from Models
- Matt Shumer declares 'Anthropic has won,' calls new model incredible — mattshumer_ · 2026-09-24
- OpenAI rolls out upgraded prompt caching and lower cached input rates for GPT-6 — rhiever · 2026-09-24
- Testing the Jeb chatbot: inconsistently biased, not neutral — calibrate it like any classifier — PawarBI · 2026-09-24
- Pokemon benchmark Paradigm 3: Astra generalizes to scrambled maps and fan-made games while rivals memorize — gleech · 2026-09-24
- AI Completes Fan-Made Pokemon Brown in 10K Steps: Real Generalization or Whack-a-Mole? — gleech · 2026-09-24
- Next-gen model names surface: Opus 5.5, Fable 5.1, GPT-6 Astra — labs said to be ~2 months ahead internally — haider1 · 2026-09-24