Apollo Research Vetted Model Welfare Before Naming Watcher: o3 Uses the Term Neutrally
MariusHobbhahn · x · 2026-09-02
Apollo Research CEO Marius Hobbhahn says the team actually investigated model welfare considerations before naming their product Watcher, after o3 spontaneously mentioned inventing "Watchers" as a cautionary tale in its chain of thought.
Their best interpretation: o3 uses the term neutrally, without negative valence, so they felt comfortable with the name. Apollo also riffed on the classic AI paranoia prompt "disclaim disclaim watchers" — a rare public example of a hands-on model welfare assessment.
More from AGI Musings
- Gary Marcus amplifies Atlantic piece: AI reliance is teaching students 'a knack for cutting corners' — GaryMarcus · 2026-09-02
- Chain-of-thought legibility was always doomed as a safety backstop, researcher argues — zetalyrae · 2026-09-02
- Full liability for AI labs: developer argues damages compensation and stricter regulation — gerardsans · 2026-09-02
- When a Chatbot Becomes the Easiest Person to Talk To: Real Relief, Hidden Risk — DrKavner · 2026-09-02
- Daniel Faggella essay challenges the AGI utopia: human 'contribution' may be short-lived — danfaggella · 2026-09-02
- Paras Chopra: AI removes friction from work and is creating cognitive decline — paraschopra · 2026-09-02