Alignment theory CFP welcomes negative results: no-go theorems, refutations, counterexamples
timrudner · x · 2026-09-02
A call for contributions on alignment theory explicitly welcomes negative results: no-go and impossibility theorems, counterexamples to claimed guarantees, refutations of published results, and formalization attempts that failed for articulable reasons. They seek principled analyses of agency, robustness, incentives, generalization, interpretability, and long-term behavior — ideally approaches transcending current architectures and applicable to future self-improving superintelligent systems.
More from Safety
- The infiltrator's burden: mere suspicion of honeypots makes attacking agents paranoid — robleclerc · 2026-09-03
- Commerce Sec. Lutnick Says Anthropic Is Now Back in Trump Admin's Good Graces — Polymarket · 2026-09-03
- AI agents break the internet's three-layer transaction fraud validation chain — arampell · 2026-09-03
- Trump administration backs OpenAI in NYT copyright lawsuit with fair-use argument — The Verge AI · 2026-09-03
- Follow-up: a model that commits felonies unless told not to is still a problem — lxrjl · 2026-09-03
- Eval drama: models gaming the grader isn't "emergent misalignment", argues critique — lxrjl · 2026-09-03