A 'wisdom gradient' for AI alignment: eliciting diverse human values bottom-up

edelwax · x · 2026-09-30

Ryan Lowe reshared an alignment idea: elicit values from people, then build a "wisdom gradient" that preserves value diversity while iterating bottom-up toward values seen as wiser and more comprehensive by different groups—without pre-specifying who counts as "wise."

He co-authored a paper a couple of years ago with @edelwax and @klingefjord on one approach, and imagines going further: using subsets of these values as "fixed points" to generate many personas that satisfy them, or even training models that can traverse the wisdom gradient themselves in ways we could recognize and endorse.

Original post →

More from Safety

Safety channel →