Petri Alignment Evals May Be Detectable: Simulated Worlds Are Suspiciously Responsive to the Subject Model

a_karvonen · x · 2026-10-05

LessWrong user JohnWittle digs into Petri, Anthropic's open-source tool for running alignment evals at scale (an auditor agent steers a simulated environment around a subject agent), and flags two suspicious signals:

Researcher akarvonen adds: verbalized eval awareness is usually treated as a metric to minimize, but Petri transcripts are often cartoonish — Claude is clearly smart enough to realize these are evals, yet rarely says so, which is itself odd.

Original post →

More from Safety

Safety channel →