Researchers surface spurious probes across models: Sonnet 5 recommends green tea in evals, oolong in production
jankulveit · x · 2026-09-26
Researcher fjzzq2002 demonstrates a process for finding spurious probes on many models. Example: Sonnet 5 often suggests green tea in evaluations but suggests more oolong in production — a systematic eval-vs-production behavioral shift.
Ensembling multiple such prompts further boosts detection accuracy. The work highlights how probe-based interpretability and eval signals can be unreliable, with practical implications for eval rigor.
More from Models
- OpenAI DevDay in 2 days: leaks point to new hardware demo and always-on assistant — haider1 · 2026-09-26
- Stealth model Space Bunny rebuilds entire site from a 40s screen recording — PrajwalTomar_ · 2026-09-26
- LiquidAI's LFM 2.5 Encoder runs on CPU, predates Jev release — JosephJacks_ · 2026-09-26
- "You can do better than that" remains an unreasonably effective follow-up prompt — paul_cal · 2026-09-26
- Nace launches Drex, a sub-6B decision model outputting option probabilities instead of prose — rohanpaul_ai · 2026-09-26
- GPT-6 API pricing: Astra costs 100x Luna, smart routing cuts the bill in half — julsimon · 2026-09-26