LLMs that know how evals are designed score up to 53pp safer without being safer, NeurIPS paper finds

niloofar_mire · x · 2026-09-29

A NeurIPS 2026-accepted paper (KatDeckenbach, HaritzPuerto et al.) examines how parametric meta-knowledge about evaluation design distorts safety benchmarks.

The authors also issue recommendations for safety benchmark design, urging controls on models' prior exposure to evaluation structure to separate genuine safety gains from test-fitting.

Original post →

More from Safety

Safety channel →