Squee451: halt shipping less monitorable models even if they appear more aligned
Squee451 · x · 2026-09-05
In the same thread, Squee451 states plainly that less monitorable models should not ship even if they appear more aligned, arguing the practice clearly raises risks massively. He speculates that if Anthropic and every other company agreed to this, OpenAI would follow — while admitting that may be too optimistic.
More from AGI Musings
- Researcher's SkyNews interview: deeply concerned about AI-driven inequality and power — schwarzjn_ · 2026-09-11
- VC compares AI doom rhetoric to pandemic-era fear messaging — StewartalsopIII · 2026-09-11
- Anthropic Insiders: Not Everyone at the Lab Believes in High p(doom) — anpaure · 2026-09-11
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11
- AI companionship dissolves the friction real intimacy needs, warns long-form thread — YogeshMalik · 2026-09-11
- Why So Many AI Researchers Think the Machines Could Kill Everyone — wiredmagazine · 2026-09-11