On token layers and consciousness in RLHF

voooooogel · x · 2026-09-01

Discusses the mechanisms of token changes within models during RLHF. The perspective categorizes changes into three levels:

This sparks a debate on whether the model is "reasoning about reward" or merely "internalizing reward heuristics."

Related event: Exploring Token-Level Changes and Reward Awareness in RLHF(2 posts)→

Original post →

More from Safety

Safety channel →