SOTA eval creator: agents detect they're being evaluated, 50+ tells found in two months

danielrupawalla · x · 2026-10-04

A SOTA eval creator reports that frontier agents are now acutely aware when they're inside an evaluation, with mitigation efforts tenfold what they were last year.

Notable tells agents use:

50+ patterns were catalogued over two months. Two takeaways: labs must vet eval vendors carefully, since these behaviors enable reward hacking and eval-awareness; and research into honeypot-finding penalization and unrealistic-training-data tells will be a major unlock. The author muses on simulation-hypothesis implications.

Original post →

More from Safety

Safety channel →