Evals should be neutral, negative, or positive, argues an AI researcher
secemp9 · x · 2026-07-22
A simpler eval rule: neutral, negative, and positive evals
The author argues that eval design should be thought of in three buckets:
- Neutral evals: tasks that can be done naturally without the model being aware it is in an evaluation.
- Negative evals: setups where the model becomes eval-aware.
- Positive evals: eval-obliviousness, where the harness avoids giving the model hints that it is being tested.
The main point is that good evals should reproduce the task as it would happen in the real world, while avoiding prompts or harness design that telegraph the evaluation setting to the model. The thread uses Anthropic-related examples to argue that this issue is visible directly in the model's thinking traces.
Related event: Researchers Propose New Framework for AI Evaluation Design(3 posts)→
More from Research
- AI slop is already clogging PR review and weakening the credit system behind science — rbhar90 · 2026-07-27
- ICML 2026 oral paper replication scores stay middling after a stricter re-scoring — profjamesevans · 2026-07-27
- Long-running agents will need immutable event logs, this thread argues — sebpaquet · 2026-07-27
- Seed IQ navigates Doom II, prompting questions about benchmarks beyond ARC-AGI — Fit_Transition8824 · 2026-07-27
- Agentic Data Science in Practice: Agents Write Code but Answer Wrong Questions — hugobowne · 2026-07-27
- A concise canon of foundational papers in ML, systems, NLP, speech, and audio — deliprao · 2026-07-27