Evals should be neutral, negative, or positive, argues an AI researcher

secemp9 · x · 2026-07-22

A simpler eval rule: neutral, negative, and positive evals

The author argues that eval design should be thought of in three buckets:

The main point is that good evals should reproduce the task as it would happen in the real world, while avoiding prompts or harness design that telegraph the evaluation setting to the model. The thread uses Anthropic-related examples to argue that this issue is visible directly in the model's thinking traces.

Related event: Researchers Propose New Framework for AI Evaluation Design(3 posts)→

Original post →

More from Research

Research channel →