Alignment theory CFP welcomes negative results: no-go theorems, refutations, counterexamples

timrudner · x · 2026-09-02

A call for contributions on alignment theory explicitly welcomes negative results: no-go and impossibility theorems, counterexamples to claimed guarantees, refutations of published results, and formalization attempts that failed for articulable reasons. They seek principled analyses of agency, robustness, incentives, generalization, interpretability, and long-term behavior — ideally approaches transcending current architectures and applicable to future self-improving superintelligent systems.

Related event: Alignment theory call seeks superintelligence-ready analysis across ten fields(4 posts)→

Original post →

More from Safety

Safety channel →