Evaluating Small Models: Multiple-Choice Tests Are Easy to Game
mervenoyann · x · 2026-07-14
AI researcher Merve Noyant points out that when evaluating large models, most multiple-choice question assessments (MCQA) actually make it easy for models to "cheat": even if a model gets stuck in a loop or hallucinates, the structure of the multiple-choice options helps it organize and anchor the correct answer. Consequently, she prefers researching small models and emphasizes the importance of conducting "vibe tests."
Related event: Multiple-Choice Evaluations May Overestimate LLMs(2 posts)→
More from Research
- OpenForecaster uses daily news to improve language-model forecasting — Cohere_Labs · 2026-07-21
- SenseTime unveils U1 Pro and open-sources a 50M-sample vision dataset at WAIC 2026 — 机器之心 · 2026-07-21
- Baseten study finds new facts in LLM weights are fragile unless trained from many restatements — alex_verem · 2026-07-21
- Kimi K3 and Fable 5 now look much closer than the old open-vs-closed gap — FinanceYF5 · 2026-07-21
- uv-scripts/ocr returns to the top of Hugging Face datasets with a JSON model picker — vanstriendaniel · 2026-07-21
- DeepSearch-World trains web agents with 420K verifiable QA tasks — HKUST · 2026-07-21