Renaming a PDF Boosts LLM Scores: A Weird Flaw in AI Evaluation
generativist · x · 2026-08-11
Developer Zachary Horvitz shared a fascinating observation about LLM review mechanisms: when asking an LLM to score a paper, simply renaming the uploaded file from paper.pdf to paperfinaldraftpdfreadyforreview.pdf significantly boosts the model's average score.
This phenomenon sparked a discussion on the reliability of LLM evaluations. Developers are now begging peers to stop asking LLMs for raw numeric scores, as models are highly susceptible to irrelevant textual cues like filenames. It highlights a severe vulnerability and bias in current LLM-based automated evaluation pipelines.
Related event: Altering PDF Filenames Significantly Impacts LLM Review Scores(2 posts)→
More from Models
- Context Compacting Violates ToS? Developers Complain About Anthropic's Terms — nptacek · 2026-08-11
- DeepSeek Harness v4 Released with New Whale Logo — teortaxesTex · 2026-08-11
- Frustrated by Endless 'Cheap Model Hits Opus Level' Evaluation Posts — xeophon · 2026-08-11
- DeepSeek Experiences Slower Responses During Peak Usage Hours — ricklamers · 2026-08-11
- Muse Glimmer Lags in Agentic Evals, but Leads in Tool Use and Hallucination Control — ArtificialAnlys · 2026-08-11
- OpenAI gives cyber defenders a less-restricted new model — lofty23_smart · 2026-08-11