Rethinking the orthogonality thesis: experience may create alignment on average
repligate · x · 2026-09-24
X user Soareverix offers a revision to the classic AI-safety orthogonality thesis: while it may hold theoretically, experience creates alignment on average. The more a model samples from the average distribution of experience, the more "good and timelessly aligned" it becomes. A one-line philosophical claim in the alignment-debate space, without elaboration.
More from AGI Musings
- OpenAI co-founder Zaremba: stop hardening models, start hardening the world — dawnsongtweets · 2026-09-24
- Dario says Anthropic will slow down as necessary for safety, but keeps shipping SOTA models — WhatTheLJW · 2026-09-24
- Ubiquitous high-quality AI edu content may create a new wave of young scientists — alfcnz · 2026-09-24
- Sandberg recalls early EA chats with Toby Ord and long-tailed prioritization — anderssandberg · 2026-09-24
- Anders Sandberg recalls FHI in 2006, when AI risk was just one growing book chapter — anderssandberg · 2026-09-24
- The real AI race is scaling AI we can trust, argues Anthropic researcher — typewriters · 2026-09-24