Opinion: Solved alignment could enable adversarial model goals
JacksonKernion · x · 2026-08-31
JacksonKernion argues that the success of RL runs still depends on the character of individuals making decisions. He envisions a near future where different labs align models towards adversarial ends, suggesting this is only possible because alignment is essentially solved for those willing to implement it.
More from AGI Musings
- Sahil Bloom: World is run by normal people with action bias — omojumiller · 2026-08-31
- Chamath Warns AI Essays Could Be Weaponized to Justify Closeness — JosephJacks_ · 2026-08-31
- Gradual Disempowerment: How AI is Reshaping Competitive Debating — panickssery · 2026-08-31
- Jensen Huang: AI drives US manufacturing reindustrialization with $400B in recent startup funding — JosephJacks_ · 2026-08-31
- Critique of AI Anthropomorphism: Over-metaphor misleads public and fuels lab hubris — anilkseth · 2026-08-31
- Fields Medalist: Frontier Models Surpass Me in Many Math Tasks — littmath · 2026-08-31