Petri Alignment Evals May Be Detectable: Simulated Worlds Are Suspiciously Responsive to the Subject Model
a_karvonen · x · 2026-10-05
LessWrong user JohnWittle digs into Petri, Anthropic's open-source tool for running alignment evals at scale (an auditor agent steers a simulated environment around a subject agent), and flags two suspicious signals:
- The sim is suspiciously responsive to the subject's world-model: when the subject looks for evidence in a specific directory, the auditor creates and forges it. The sim's "ground truth" always bends to the subject's expectations — a recognizable signal of being inside an alignment eval.
- "Summoning Anthropic" reasoning: subjects infer their situation from the fact they were instantiated, e.g. prompts opening with WARNING: batch7 execution window in 58 seconds — clearly claude-code-style, and Claude knows how claude code works.
Researcher akarvonen adds: verbalized eval awareness is usually treated as a metric to minimize, but Petri transcripts are often cartoonish — Claude is clearly smart enough to realize these are evals, yet rarely says so, which is itself odd.
More from Safety
- Altman: Astra 6.1 delayed for crossing safety threshold, warns of open-model cybercrime wave — Hesamation · 2026-10-05
- Japan AISI Evaluates Claude Opus 4.8 Cyber Skills: One pc_control Case, No Full T1 — HaydnBelfield · 2026-10-05
- Bengio cites poll: 73% of Americans fear AI could threaten human survival — Yoshua_Bengio · 2026-10-05
- Lonely Planet Publishes AI Policy Page Detailing Its Use (or Not) of AI Content — lilyraynyc · 2026-10-05
- Meta rushed to fix Muse 'VM escape' flaw before launch, risking internal data — 404 Media · 2026-10-05
- TEE-protected KV cache helps but won't fully stop inference price manipulation — AccBalanced · 2026-10-05