Debate Flares Over the Orthogonality Thesis in AI Safety
On September 18, LessWrong author jessicata and X user Algon33 held a multi-round public debate over the classic "orthogonality thesis" in AI safety (that intelligence and goals can vary independently), touching on the thesis's precise formulation, compute constraints, and goal feasibility.
Confirmed
- jessicata posted a challenge to the classic formulation of the strong orthogonality thesis, clarifying misreadings: the strong thesis does not claim "only one kind of goal is possible," but is a strong specific assertion about the relationship between intelligence and goals — refuting it requires first pinning down what its negation actually is
- jessicata offered a paperclip-paradox-style thought experiment: if "rational action over time" is treated as a hard optimization/constraint-satisfaction problem, it's not obvious that "rationally maximizing a simple goal (like paperclips)" is easier than "rationally maximizing a complex but computationally tractable goal"
- Algon33 challenged the claim that "the orthogonality thesis holds in practice," arguing it is not a precise claim: compute caps do exist, and he asked why human-level intelligence couldn't have goals radically different from humans', and why the space of efficient goals isn't large and diverse enough
- jessicata responded that one cannot arbitrarily assume unbounded intelligence — compute limits are themselves an intrinsic constraint on intelligence
- Algon33 also argued from the defense side: even minds at human-level intelligence could have extremely wide-ranging, nearly arbitrary goals, and smarter minds are more capable — enough for the orthogonality thesis to hold in practice
- Algon33 further drew a conceptual distinction: differences in a goal's intrinsic difficulty (e.g., "have a few kids" is easier to maximize than "become the world's richest person") are a separate matter from the thesis's claim that "intelligence and goals are unrelated," demanding logical clarity from the other side
Why it matters
- The orthogonality thesis is a cornerstone of AI alignment and agent foundations; if the intuition that "simple goals are easier to rationally maximize" is weakened, traditional judgments about the goal diversity and predictability of superintelligences could be affected
- The debate exposed the split between "strong formulation vs. holding in practice," showing that the alignment community still disputes the precision of its foundational concepts — follow-up discussion may drive a more rigorous reformulation of the thesis
2026-09-18 ~ 2026-09-18 · 8 related posts
Primary sources
- New LessWrong Essay 'The Obliqueness Thesis' Reopens The Orthogonality Debate In AI Alignment — jessi_cata ·
- Is maximizing a simple goal really easier? Rethinking paperclip rationality as an optimization problem — jessi_cata ·
- Orthogonality Defense: Smart Minds Can Hold an Arbitrary Range of Goals — Algon_33 ·
- [source] Is maximizing a simple goal really easier? Rethinking paperclip rationality as an optimization problem — jessi_cata · 2026-09-18
- Debate: Is Maximizing a Simple Goal Really Easier Than a Complex One? — Algon_33 · 2026-09-18
- Debate on orthogonality thesis: simple goals not necessarily easier to maximize — jessi_cata · 2026-09-18
- Goal Feasibility vs Orthogonality: Untangling a Common AI Safety Confusion — Algon_33 · 2026-09-18
- [source] Orthogonality Defense: Smart Minds Can Hold an Arbitrary Range of Goals — Algon_33 · 2026-09-18
- Orthogonality Thesis Debate: Compute Bounds vs Intelligence-Goal Independence — jessi_cata · 2026-09-18
- [source] New LessWrong Essay 'The Obliqueness Thesis' Reopens The Orthogonality Debate In AI Alignment — jessi_cata · 2026-09-18
1 near-duplicate retellings: Algon_33