Can Models Detect Modified Activations?

belindazli · x · 2026-07-09

The post mentions that several recent papers have found frontier models leave very faint traces when their "activations are modified." The author and collaborators further proposed and verified that making this detection capability more robust is surprisingly easy to achieve.

Original post →

More from Research

Research channel →