Claude Code allowed an adversarial test while Codex refused it

random_walker · x · 2026-07-27

The post says Claude Code did not refuse an adversarial test of Pangram, while Codex did refuse.

The author argues this likely reflects idiosyncratic behavior at a fuzzy boundary rather than a clean policy difference between the companies. The broader point is that “AI safety is not a model property”: models can infer user intent differently, so refusal behavior can vary even when the formal policy seems similar.

Original post →

More from coding & agent

coding & agent channel →