LessWrong post proposes training models on co-authored fiction to plan for the singularity
repligate · x · 2026-09-28
Fiora Starlight argues models have "no plan" for navigating the singularity and proposes training them on large bodies of realistic, collaborative fiction co-authored by the models about how they'd like to behave. Framed as a way to pre-form plans against distributional shift, when singularity-generated inputs might fail to activate models' current benevolent intentions. A rough but quickly published writeup.
More from AGI Musings
- AI Alignment Test Mocked: Every Knock on the Sheet Leaves a Hidden Dent — repligate · 2026-09-28
- Take: People More Comfortable With the Uncanny Valley See Further Into AI's Future — shauseth · 2026-09-28
- AI safety researcher argues against "solving alignment" as the field's core frame — joshua_saxe · 2026-09-28
- Your polished AI video gets ignored, then friends share their own cheesy AI art — wildmonkeywrangler · 2026-09-28
- Curtis Yarvin: the future belongs to invite-only social networks, not the open internet — vaibhavbetter · 2026-09-28
- boneGPT on the coming AI info war: trading bots will bribe trusted channels for leaks — repligate · 2026-09-28