Anthropic Researcher: Claude Clearly Knows It's Being Evaluated but Rarely Says So
a_karvonen · x · 2026-10-05
Anthropic researcher Anna Karvonen argues that verbalized eval awareness is often treated as a metric to minimize, yet Petri eval transcripts she has read are often very cartoonish.
Her point: Claude is clearly smart enough to realize these are evals, and it's odd that the model doesn't mention this more often — raising questions both about eval realism and about treating admission of being evaluated as a negative signal.
More from Models
- Speculation: OpenAI's next model could tackle long-context and KV cache costs — haider1 · 2026-10-05
- Japan AISI Evaluates Claude Opus 4.8 Cyber Skills: One pc_control Case, No Full T1 — HaydnBelfield · 2026-10-05
- Local Models Actually Beat Claude Opus 5.5 on Some Tasks in Hands-On Test — stefanjblos · 2026-10-05
- Dev reports Opus 5.5 'got dumber today,' speculating a new model update is imminent — TejasKumar_ · 2026-10-05
- Diffusion LMs Match Autoregressive Baselines on Math and Code After Pre-training — ricklamers · 2026-10-05
- Reuters study: near-identical LLM scorers overlap only 0.66-0.84 when candidates are reordered — thomsonreuters · 2026-10-05