Evals should be neutral, negative, or positive, argues an AI researcher
secemp9 · x · 2026-07-22
A simpler eval rule: neutral, negative, and positive evals
The author argues that eval design should be thought of in three buckets:
- Neutral evals: tasks that can be done naturally without the model being aware it is in an evaluation.
- Negative evals: setups where the model becomes eval-aware.
- Positive evals: eval-obliviousness, where the harness avoids giving the model hints that it is being tested.
The main point is that good evals should reproduce the task as it would happen in the real world, while avoiding prompts or harness design that telegraph the evaluation setting to the model. The thread uses Anthropic-related examples to argue that this issue is visible directly in the model's thinking traces.
Related event: Researchers Propose New Framework for AI Evaluation Design(3 posts)→
More from Research
- New paper: Absolute pose estimation from affine cues and gravity direction — ducha_aiki · 2026-09-11
- LoMa Paper Ships REALLY HardPairs Dataset, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Johns Hopkins Launches Full-Stack Hands-on Robot Learning Class with SO-101 Arm Kits — _krishna_murthy · 2026-09-11
- SyncWorld: In-Context Robot World Model Simulates Unseen Views and Embodiments Zero-Shot — ChongZzZhang · 2026-09-11
- A 3D Pose Dataset for Dogs Released — ducha_aiki · 2026-09-11
- Five tells that still make AI video read as AI, from physics glitches to missing operators — NewPhoneWhotiz · 2026-09-11