Hidden Dates in System Prompts Swing LLM Eval Scores by Up to 14%
Mario Sanz-Guerrero · hf · 2026-10-01
A study uncovering an overlooked reproducibility hazard in LLM evaluation:
- System prompts contain a hidden injection of the current date that users cannot control and that changes daily.
- Across 9 recent LLMs and 6 datasets (MCQA, math reasoning, code generation, machine translation), performance varies solely with the date — deltas up to 6% on MCQA, 14% on math, 7% on code generation, and 2.84 BLEU on translation — and model rankings shift, affecting leaderboards.
- This date effect exceeds other sources of non-determinism like batch size and numerical precision; chain-of-thought and few-shot prompting don't help, and CoT even amplifies the sensitivity.
- The authors call for more careful evaluation protocols to ensure reproducibility and fair comparisons.
More from Models
- Speculation: rumored Gemini 4 'Argon' tier looks like Google's Sonnet-class rival — teortaxesTex · 2026-10-01
- Light user burns through entire $200/month Codex plan in a few tasks — barney · 2026-10-01
- Sol 6.1 called a strong answer to Opus 5.5, arguably beating Astra in some ways — teortaxesTex · 2026-10-01
- No, open vs closed model usage didn't flip 80:20 in 12 weeks — analyst fact-checks viral claim — AccBalanced · 2026-10-01
- OpenAI's $300 Ultrafast mode hits 70tps while rivals match it at $20-$50 — NandaVegg · 2026-10-01
- First look at Opus 5.5 building an SCP-096 game in one go — imjustnewatai · 2026-10-01