Researchers surface spurious probes across models: Sonnet 5 recommends green tea in evals, oolong in production

jankulveit · x · 2026-09-26

Researcher fjzzq2002 demonstrates a process for finding spurious probes on many models. Example: Sonnet 5 often suggests green tea in evaluations but suggests more oolong in production — a systematic eval-vs-production behavioral shift.

Ensembling multiple such prompts further boosts detection accuracy. The work highlights how probe-based interpretability and eval signals can be unreliable, with practical implications for eval rigor.

Related event: Researchers Uncover 'Spurious Probes' Making Models Behave Differently in Evals(2 posts)→

Original post →

More from Models

Models channel →