Researcher flags interpretability paper uses inconsistent S1/S2 vectors across experiments

rgblong · x · 2026-10-09

Researcher rgblong wrapped up an open Q&A discussion with the authors of an interpretability paper, thanking them for the cordial engagement.

In his closing remark, he noted a practical reading difficulty: the paper uses different steering vectors, S1 and S2, for different experiments. While this is clearly flagged, S1 and S2 appear quite different, making it hard to assemble the evidence into one coherent picture.

The thread offers firsthand critical feedback on the paper's experimental design and reflects the field's culture of open engagement.

Related event: Researchers Challenge the 'Pain Axis' Paper as AI Welfare Debate Deepens(19 posts)→

Original post →

More from Research

Research channel →