AI2D falls from 77.6% to 40.5% when multiple choice becomes free response

DatBench: Discriminative, Faithful, and Efficient VLM Evaluations

DatologyAI, :, Siddharth Joshi, Haoli Yin, Rishabh Adiga, Ricardo Monti, Aldo Carranza, Alex Fang, Alvin Deng, Amro Abbas, Brett Larsen, Cody Blakeney, Darren Teh, David Schwab, Fan Pan, Haakon Mongstad, Jack Urbanek, Jason Lee, Jason Telanoff, Josh Wills, Kaleigh Mentzer, Luke Merrick, Parth Doshi, Paul Burstein, Pratyush Maini, Scott Loftin, Spandan Das, Tony Jiang, Vineeth Dorna, Zhengping Wang, Bogdan Gaza, Ari Morcos, Matthew Leavitt

cs.LG, cs.AI

2026-01-06

DatologyAI cleans 33 VLM benchmarks. AI2D drops from 77.6% MCQ to 40.5% free response; a curated subset closely matches discrimination at 13× speedup (up to 50×).

What problem this solves

Public vision-language model (VLM) benchmarks are noisy in two different ways. Multiple choice carries a 1/N guessing floor, and the options themselves leak linguistic shortcuts. On the General split, more than 70% of items are solvable from text alone; VQA-v2 is in the same range. The bill is also large. During OLMo3 post-training, nearly 20% of the compute budget went to evaluation. High-resolution visual tokens plus a reasoning trace can run to tens of thousands of tokens per item. OCR suites often exceed 100,000 examples, and a thinking model that writes past 32K tokens can force more than 3 billion generated tokens for a single capability.

A gain of a few points on that noise floor is hard to read as a real capability change. DatologyAI treats the fix as curation of benchmarks that already exist. Thirty-three datasets are kept, then judged on three tests: the item must actually need the image and look like downstream use; stronger models must separate from weaker ones; and the score has to be worth the tokens.

Method

The pool is split into nine capabilities: charts, documents, scene text, math and logic, spatial reasoning, grounding (pointing at a region from a phrase), counting, diagrams and tables, and general visual question answering. Filtering and scoring use 27 models from 1B to 10B parameters, including Qwen3-VL, several InternVL generations, GLM-4.1V, R-4B, and Gemma-3. Generation is capped at 4,096 tokens.

Four passes touch the data only.

Where a multiple-choice item can stand alone, the options are stripped and the model must write the answer. Qwen3-30B judges semantic match, so a near-miss string is not an automatic fail. Items that only make sense as a choice among listed options stay multiple choice, but under circular evaluation: the option order is rotated, and the model scores only if it is right on every pass.

Blind items are found by deleting the image. All 27 models see text only. On open-ended sets the threshold is τ = 1: if any model is correct without the image, the item goes. Multiple choice and counting use a looser threshold, because a short answer list is easy to hit by chance. The discarded items are mostly world knowledge (a mosquito has four life stages), visual stereotypes (toilets are usually white), and pure arithmetic that never needed a figure (how often the digit 2 appears from 1 to 30).

Label noise is a two-stage cut. Items that every 1B–10B model misses are flagged, then GPT-5.2 checks them with the official answer in view. Ambiguous wording, wrong labels, and images too coarse to identify the target are removed. The pool is large enough that the policy is deliberately strict.

Discrimination is the point-biserial correlation rpb, which needs no fitted hyperparameters. It asks whether getting an item right moves with a model's overall score. Items that strong models hit and weak models miss are kept first. Negative rpb, where weaker models win, is pushed to the back. Classical item-response theory needs far more than 27 systems to fit a stable difficulty and discrimination, so rpb stands in. Up to 20% of each DatBench slice is then reserved, by hand, for frontier items the judge accepts but none of these models solve, so the subset does not saturate immediately.

DatBench is the subset after all four passes, meant for training loops. DatBench-Full stops after the quality filters. On counting and spatial reasoning the two are similar in size; elsewhere Full can be up to about 50 times larger.

Results

Changing the format is enough to move the score. On AI2D the 27 models average 77.56% in multiple choice and 40.53% in free response. The strongest multiple-choice model loses nearly 35 points. Circular evaluation has a slope steeper than 1 against vanilla multiple choice. The vanilla format often hands models a 20–30% floor; consistent accuracy for those same models stays under 50%.

ChangeBaselineResult
AI2D as free response77.56% multiple choice40.53% mean; strongest model nearly −35 points
General, blind items removedOriginal scores packed into 65–80%72.07% of text-solvable items removed; scores spread across 10–65%, about 4× the range
Spatial quality filterOriginal items42.07% dropped for ambiguity or resolution; ChartQA Pro 17.2%, MMMU-Pro 24.3%; grounding and General under 1%
Keep 40% by rpbRandom keep at the same rateAbout 90% of total discrimination; random keeps less than half that signal
Released DatBenchFiltered full set13× faster on average, up to about 50×

Pearson correlations split the suite. Charts and general VQA sit at r = 0.90, math and general VQA at r = 0.76. Grounding runs against document understanding at r = −0.29 and against OCR at r = −0.19. Math against spatial reasoning is also r = −0.19. Clustering puts charts, math, and general VQA in one group, and OCR, spatial reasoning, and diagrams in the other.

GLM-4.1V-9B is a perception specialist: 66.4% on diagrams, 36.8% on spatial reasoning, 17.4% on math. Qwen3-VL-4B is unusually even, at 71.0% documents, 77.9% OCR, and 59.9% reasoning. R-4B posts the highest math score, 43.4%, and the lowest spatial score, 11.4%.

Thinking models, which write a reasoning trace before the answer, are compared with the instruct twin from the same family. The relative gain is (thinking accuracy − instruct accuracy) / instruct accuracy. Math is about +36.8% and charts about +10.8%. OCR is about −53.5% and document understanding about −47.8%. A correct thinking trace is about 425 tokens; a wrong one is about 1196.9 tokens, roughly 14× the token use of non-thinking models. On perception items the extra tokens are mostly a loop that does not finish.

The vision gap, multimodal score minus text-only score, is 60.2% for counting and 42.3% for grounding, but only 13.0% for math and 14.9% for spatial reasoning. A large slice of the math number is the language model guessing.

Why it matters

Use DatBench inside a training loop and DatBench-Full for the final report. The 13× figure is an average over nine capabilities. Counting and spatial reasoning are already small after the filters, so the speedup there is modest. Charts and grounding sit near the diagonal against the original benchmarks, so the ranking survived. General VQA and documents get a steeper slope, and models that used to share a narrow band come apart.

Extra tokens pay on reasoning tasks and cost accuracy on OCR and documents, and the failures are the expensive ones. A product path that is mostly recognition should not turn thinking on by default.

Lower scores here come from deleting guesses, blind items, and bad labels. The tasks were not rewritten into a harder definition.

Limitations

rpb is estimated on 1B–10B models with a 4,096-token cap. The paper states that the high-discrimination items will shift once models are larger or traces are longer. There is no evidence this subset still separates frontier systems.

Selection maximizes per-item correlation and does not constrain diversity. A cluster of near-duplicate items can all look highly discriminative and fill the subset. The pipeline relies on the source mix and does not test redundancy.

Sending every unanimous failure to GPT-5.2 puts genuine hard items and broken labels in one bin. The paper admits the judge can call a valid specialist argument a defect, and the filter is conservative on purpose. After 42.07% of spatial items are removed, nothing in the paper checks, with human relabeling, whether the remainder still measures spatial understanding.

Rank correlation saturates quickly here because 1B and 8B models are far apart, so a random subset can preserve a coarse order. rpb is still fit on these same 27 models. Generalization to unseen architectures is asserted, not measured on a held-out set. Agreement between the Qwen3-30B answer judge and the GPT-5.2 label judge is also unreported.

Terms

Source

What people are saying

Related papers

All paper explainers