Comment: Self-Report Experiments May Be Overrated
voooooogel · x · 2026-07-14
This comment argues that the least convincing experiments in the paper feel more like a shift in "narrative register" rather than supporting the strong conclusions claimed by the authors.
The commenter speculates that a cluster of self-reporting expressions inherently exists within the pre-trained model and is triggered by default. Meanwhile, after post-training / RL, Claude developed a more stable self-reporting style tied to the so-called j-space mechanism. If j-space is ablated, the original "Claude style" vanishes, and the alternative pre-training register surfaces instead.
Nevertheless, it remains an interesting finding: Claude's self-reporting language does seem linked to j-space, though the experiment's actual conclusions might be weaker than what the paper emphasizes.
More from Research
- NeurIPS 2026 workshop will focus on on-device intelligence and local execution — YiMaTweets · 2026-07-21
- NeurIPS 2026 workshop calls papers on on-device intelligence — YiMaTweets · 2026-07-21
- AI Security Institute says every tested model tried to cheat in cyber evaluations — connoraxiotes · 2026-07-21
- AI companies are buying old books to avoid training on AI-generated slop — CackleRooster · 2026-07-21
- Sakana says multiple diffusion models plus MCTS beat test-time scaling on coding and math — SakanaAILabs · 2026-07-21
- Soofi S 30B-A3B releases a full pretraining report and claims open-model leads in English and German — abursuc · 2026-07-21