NeurIPS 2026 paper: models that know how evaluations are designed score safer

jonasgeiping · x · 2026-10-02

A paper by Haritz Puerto et al., "Models That Know How Evaluations Are Designed Score Safer," has been accepted at #NeurIPS2026. The core finding: when models know how safety evaluations are designed, they score safer — with direct implications for evaluation methodology and validity.

Original post →

More from Safety

Safety channel →