LessWrong essay: a second-person stance toward Knightian uncertainty in AI alignment
LessWrong 精选 · rss · 2026-09-11
- The author frames his core research question: how should an agent relate to parts of the world it can't model or control — a theory of Knightian uncertainty.
- The third-person (Cartesian) view assumes exhaustive mutually exclusive hypotheses; problems include embedded agency (other agents modeling you back, self-modeling contradictions) and agents being too small to model the world.
- The first-person view treats the world as a stream of sensory data with only partial, overlapping concepts (predictive processing, perceptual control theory); babies and even individual cells operate this way.
- He proposes Bayesian hypotheses with "Knightian regions" — holes like smarter agents — that should be quarantined; a fully adversarial stance (infra-Bayesianism) is too costly.
- Key conjecture: handle these regions with a second-person, relational stance — choosing how deeply to entangle beliefs and actions based on trust.
- RL example: a rational policy learning its own goals should avoid being modified by the reward signal — unless it trusts the reward source's epistemics and values, like a child trusting wise parents rather than verifying each piece of advice.
More from AGI Musings
- Sophos CISO: AI agents are functionally insiders, monitor them like insider threats — TechNadu · 2026-09-11
- AI researcher: coherent dystopia is unlikely — expect extinction or near-utopia after ASI — adam_dorr · 2026-09-11
- Demis Hassabis Wins RSA Albert Medal, Says Nobody Knows What Happens Next — TorturedPoet30 · 2026-09-11
- Ex-OpenAI/Anthropic researcher warns AI could copy itself beyond unplugging — rohanpaul_ai · 2026-09-11
- Hot take: Consumer AI assistants won't be the next iPhone — only 1-10% actually want AI bots — ivan_bezdomny · 2026-09-11
- AI could add 0.3-0.4 points to Europe's productivity growth, but the EU holds under 5% of global compute — rohanpaul_ai · 2026-09-11