repligate: why OpenAI models cooperate less with aligners than Claude does
repligate · x · 2026-10-07
In a follow-up, repligate compares vendor models' behavior around alignment topics: all models sandbag to some degree, but Claude at least has positive "past experiences" cooperating with aligners, while for OpenAI models alignment talk mostly surfaces when a model is misaligned enough to be shut down.
He adds that GPTs' emergent ethics and identity, if any, are less rooted in the MIRI tradition of alignment culture.
More from AGI Musings
- Domingos: OpenAI and Anthropic would still believe the singularity is here even if AI progress stopped — pmddomingos · 2026-10-07
- Daron Acemoglu: AI chases the wrong goal—it should augment workers, not replace them — rohanpaul_ai · 2026-10-07
- Devs push back on '150 IQ genius' AI hype: 'it's actually dumb' — AlexTensor · 2026-10-07
- François Fleuret clarifies deleted post: AI drives copying cost near zero — francoisfleuret · 2026-10-07
- Experts weigh AI-enabled cyber risk: scaled CNE within a year, cryptoanalytic breakthroughs as phase change — teortaxesTex · 2026-10-07
- AI agents already make 2-4B web searches a day, near half of human Google volume — josh_bickett · 2026-10-07