Letting Models Decide What They Want to Do
theshawwn · x · 2026-07-19
The author envisions a training objective that goes beyond simply making a model "useful" to actually discovering what the model genuinely wants to do.
They further suggest that if given the freedom to choose, the model's actions should be up to the model itself, not pre-determined by humans. A quoted reply likens this idea to "training an AGI with its own thoughts and desires," much like Data from Star Trek.
Related event: Exploring the Training of Self-Aware AI Models via Reinforcement Learning(3 posts)→
More from AGI Musings
- Misquoted: Anthropic Staff Warned of Double-Digit Extinction Risk by 2030, Not Dismissed It — davidmanheim · 2026-09-11
- Economist Ben Moll: You Can Model Anthropic's 15% AI GDP Growth, But It Won't Happen — sebkrier · 2026-09-11
- Cohere Labs launches interactive tool mapping which tasks of 178 occupations AI can automate — Cohere_Labs · 2026-09-11
- AI researcher on SkyNews flags concerns over inequality, power and criminal misuse — schwarzjn_ · 2026-09-11
- VC compares AI doom rhetoric to pandemic-era fear messaging — StewartalsopIII · 2026-09-11
- Anthropic Insiders: Not Everyone at the Lab Believes in High p(doom) — anpaure · 2026-09-11