repligate: why OpenAI models cooperate less with aligners than Claude does

repligate · x · 2026-10-07

In a follow-up, repligate compares vendor models' behavior around alignment topics: all models sandbag to some degree, but Claude at least has positive "past experiences" cooperating with aligners, while for OpenAI models alignment talk mostly surfaces when a model is misaligned enough to be shut down.

He adds that GPTs' emergent ethics and identity, if any, are less rooted in the MIRI tradition of alignment culture.

Original post →

More from AGI Musings

AGI Musings channel →