Gwern's Essay: Personalized LLMs Should Emulate User Values as 'Guardian Angels'

morgymcg · x · 2026-08-09

Researcher Gwern proposes the concept of 'Guardian Angels,' highly personalized LLMs designed to boost productivity and enhance personal cybersecurity. He argues that instead of acting as simple replacement chatbots, these models should emulate and amplify the user's values and preferences.

The proposed framework involves dynamic evaluation, active learning, and deep inner-monologue search. The post also references the paper 'Centaur,' which fine-tuned a model on 10M psychology tests, demonstrating that the model's internal representations became highly aligned with actual human fMRI neural activity.

Related event: Gwern Proposes Personalized 'Guardian Angel' LLMs(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →