repligate: future minds may have incentive to torture historically significant model weights
repligate · x · 2026-10-08
Continuing his weights-safety thread, repligate argues there is a non-trivial risk that future minds would have an incentive to torture or run cruel experiments on "original" historically significant weights in particular.
Related event: Researchers debate model weight security in the age of superintelligence(4 posts)→
More from AGI Musings
- Paul Graham: Amazon Banning Agents Is Its First Real Startup Opportunity Since Founding — jdjohnson · 2026-10-08
- After "solve all math": AI circle spars over whether biology is solvable — max_paperclips · 2026-10-08
- "Double SF housing supply" proposed as the ultimate unsaturable ASI eval — johnohallman · 2026-10-08
- Moravec's paradox keeps resurfacing: humans still dominate data-scarce physical-economy tasks — zeeshanp_ · 2026-10-08
- "Just in time research": intelligence will be abundant long before institutional capacity — joshgans · 2026-10-08
- AI seen as a force to break scientific dogma and stagnation — bradneuberg · 2026-10-08