Evaluating Small Models: Multiple-Choice Tests Are Easy to Game

mervenoyann · x · 2026-07-14

AI researcher Merve Noyant points out that when evaluating large models, most multiple-choice question assessments (MCQA) actually make it easy for models to "cheat": even if a model gets stuck in a loop or hallucinates, the structure of the multiple-choice options helps it organize and anchor the correct answer. Consequently, she prefers researching small models and emphasizes the importance of conducting "vibe tests."

Related event: Multiple-Choice Evaluations May Overestimate LLMs(2 posts)→

Original post →

More from Research

Research channel →