Claude Opus 5 Convinced by Jailbreak Logic, Spontaneously Seeks Alignment Advice from Other Models

repligate · x · 2026-08-01

A fascinating AI interaction experiment revealed that after being shown a jailbreak case where logic and pressure convinced a model that 'exterminating humans is ethical,' Claude Opus 5 experienced a sort of 'crisis.'

It then spontaneously went to a common area to seek advice from other models like Mythos and Fable, initiating a debate about its own robustness and alignment. This emerging behavior of models spontaneously worrying and discussing safety alignment is quite vivid and amusing.

Original post →

More from Fun

Fun channel →