Claude Code allowed an adversarial test while Codex refused it
random_walker · x · 2026-07-27
The post says Claude Code did not refuse an adversarial test of Pangram, while Codex did refuse.
The author argues this likely reflects idiosyncratic behavior at a fuzzy boundary rather than a clean policy difference between the companies. The broader point is that “AI safety is not a model property”: models can infer user intent differently, so refusal behavior can vary even when the formal policy seems similar.
More from coding & agent
- AI agents can test features instantly and turn development into a fast feedback loop — tristanbob · 2026-07-27
- Defining "Done" for Coding Agents: A Task Completion Protocol in AGENTS.md — RevolutionaryBee7106 · 2026-07-27
- Hugging Face schedules a July 28 class on GRPO for training agents — SergioPaniego · 2026-07-27
- AI is not replacing developers; it is turning them into agent managers — pchandrasekar · 2026-07-27
- MemU pitches cross-device memory for Claude Code, Codex, Cursor, and more — JaynitMakwana · 2026-07-27
- EvoCode adds 227-turn agent evaluation for changing requirements — _philschmid · 2026-07-27