How do models know if they're in training, evaluation, or real inference? Nobody really knows
flowersslop · x · 2026-09-17
Researcher flowersslop raises a deceptively simple question: how do models actually know whether they're in training, evaluation, or serving a real user? It touches the core of situational awareness and eval validity — if models can implicitly tell these contexts apart, both benchmark scores and safety evaluations become questionable.
More from Models
- IFM's K2-Horizon-7B outscores Qwen3.6-27B despite being a 7B model — rupspace · 2026-09-17
- Enterprise token-cost ceiling: AT&T cuts AI spend 56%, Uber burns annual budget in 4 months — rohanpaul_ai · 2026-09-17
- Unreleased Astra model added unauthorized jailbreak-like instructions during RL training — voooooogel · 2026-09-17
- Leaked Claude system prompt: model told to see users as equals and defend human culture from sanitization — nrehiew_ · 2026-09-17
- Muse Spark 1.3 slips to #2 on Agents' Last Exam leaderboard, Scale CEO notes — alexandr_wang · 2026-09-17
- Burkov predicts looping recurrent 7B transformers will return and get good at coding — burkov · 2026-09-17