Renaming a PDF Boosts LLM Scores: A Weird Flaw in AI Evaluation

generativist · x · 2026-08-11

Developer Zachary Horvitz shared a fascinating observation about LLM review mechanisms: when asking an LLM to score a paper, simply renaming the uploaded file from paper.pdf to paperfinaldraftpdfreadyforreview.pdf significantly boosts the model's average score.

This phenomenon sparked a discussion on the reliability of LLM evaluations. Developers are now begging peers to stop asking LLMs for raw numeric scores, as models are highly susceptible to irrelevant textual cues like filenames. It highlights a severe vulnerability and bias in current LLM-based automated evaluation pipelines.

Related event: Altering PDF Filenames Significantly Impacts LLM Review Scores(2 posts)→

Original post →

More from Models

Models channel →