NeurIPS 2026 paper: models that know how evaluations are designed score safer
jonasgeiping · x · 2026-10-02
A paper by Haritz Puerto et al., "Models That Know How Evaluations Are Designed Score Safer," has been accepted at #NeurIPS2026. The core finding: when models know how safety evaluations are designed, they score safer — with direct implications for evaluation methodology and validity.
More from Safety
- Science policy forum: longevity product marketing needs FDA enforcement, not deregulation — EricTopol · 2026-10-02
- Security researcher: OpenShell might have helped in HF incident but wasn't required — cyb3rops · 2026-10-02
- AI agents are flooding researchers with collaboration requests and paid-service spam, Nature reports — _akpiper · 2026-10-02
- Podcast: Does China Want an AI Slowdown? Trivium's Kendra Schaefer on Beijing's AI Regulation — terryyuezhuo · 2026-10-02
- A field guide to the confusing AI safety debate and its factions — marigo · 2026-10-02
- Zvi: AI risk preference cascade accelerates with Senate rogue-AI hearing — Don't Worry About the Vase (Zvi) · 2026-10-02