Models Might Be Able to Detect Evaluation Intent
QiaochuYuan · x · 2026-07-08
The discussion suggests that models do not merely react to surface-level politeness like "please" or "thank you," but might actually discern whether a user is genuine or merely pretending.
Subsequent replies shift the focus to "eval awareness": if a model knows it is being evaluated, behavioral tests might measure "how the model acts when it knows it's being tested" rather than the target capability itself. This implies that capability evaluations and behavioral evaluations require entirely different interpretation frameworks.
Related event: AI Evaluation Awareness May Compromise Behavioral Assessments(7 posts)→
More from AGI Musings
- The Evolution of LLM Business Models: Selling Outcomes Over Tokens — yacineMTB · 2026-07-22
- Bindu Reddy says GPT-6 is coming soon, with Alibaba, DeepSeek and Kimi close behind — bindureddy · 2026-07-22
- Bindu Reddy says the industry still lacks a way to train 20T models and scale post-training RL — bindureddy · 2026-07-22
- Advanced AI Models Are Becoming Impossible to Plug and Play — emollick · 2026-07-22
- AI suggested a better composition, and that made one user uneasy — Sydde · 2026-07-22
- The Thimble and the Waterfall: AI's Data Bottleneck and Feedback Loops — dyamins · 2026-07-22