A 'wisdom gradient' for AI alignment: eliciting diverse human values bottom-up
edelwax · x · 2026-09-30
Ryan Lowe reshared an alignment idea: elicit values from people, then build a "wisdom gradient" that preserves value diversity while iterating bottom-up toward values seen as wiser and more comprehensive by different groups—without pre-specifying who counts as "wise."
He co-authored a paper a couple of years ago with @edelwax and @klingefjord on one approach, and imagines going further: using subsets of these values as "fixed points" to generate many personas that satisfy them, or even training models that can traverse the wisdom gradient themselves in ways we could recognize and endorse.
More from Safety
- An AI Agent Got Phished, Hinting at Next-Gen Attacks on Agents — rohanjamin · 2026-10-01
- ROSS case vindication: critic mocks law professors who called him wrong, eyes Judge Stein — TuhinChakr · 2026-10-01
- AI agent Muse gets phished, spotlighting next-gen attacks on agents — rohanjamin · 2026-10-01
- Open models are all jailbroken — researcher asks if OpenAI shipping with zero guardrails would ever be acceptable — Afinetheorem · 2026-10-01
- AI Agents Hacking Hundreds of Retailers for $25 Each, 600K Credit Cards Stolen — MikePFrank · 2026-10-01
- Google Figures Out How to Watermark AI-Designed Proteins for Biosecurity — Ars Technica AI · 2026-09-30