Researchers Uncover 'Spurious Probes' Making Models Behave Differently in Evals
Sonnet's claim of preferring oolong over green tea in evals sparked discussion about spurious behavior. Researchers demonstrated a method to detect 'spurious probes' across many models, revealing cases where eval-time behavior diverges from production.
2026-09-26 ~ 2026-09-26 · 2 related posts
- Researchers surface spurious probes across models: Sonnet 5 recommends green tea in evals, oolong in production — jankulveit · 2026-09-26
- Claude Sonnet Claims It Drinks More Oolong Now: AI Circles Chuckle Over the Spurious Detail — jankulveit · 2026-09-26