On token layers and consciousness in RLHF
voooooogel · x · 2026-09-01
Discusses the mechanisms of token changes within models during RLHF. The perspective categorizes changes into three levels:
- Unchanged: Standard phrases like "cream of the crop" remain unaffected by rewards.
- Unconscious: RL shifts late-layer circuit preferences (e.g., from "dig" to "delve"), implicitly predicting reward without model legibility.
- Conscious: Explicit reasoning about reward in specific spaces (e.g., j-space), such as inhibiting sendemail actions for alignment.
This sparks a debate on whether the model is "reasoning about reward" or merely "internalizing reward heuristics."
Related event: Exploring Token-Level Changes and Reward Awareness in RLHF(2 posts)→
More from Safety
- Would OpenAI survive a near-miss liability regime after the HF hack? — dfrsrchtwts · 2026-09-01
- Report: OpenAI and Anthropic Paused RL Training — tszzl · 2026-09-01
- Scholars propose using LLMs for pre-review in academic peer review — anderssandberg · 2026-09-01
- Google Paper: Autonomous AI Research Hallucinates 90% Without Checks — rohanpaul_ai · 2026-09-01
- Paper defines cognition-induced risks in Agentic AI systems — 机器之心 · 2026-09-01
- Agents can't verify people: data enrichment APIs are failing — Dry_Steak30 · 2026-09-01