Anthropic evals show GLM-5.3 achieves control-flow hijacks in 4% of trials, crossing a threshold

teortaxesTex · x · 2026-09-30

Citing Anthropic's internal evals, GLM-5.3 develops full control-flow hijacks in 4% of agentic safety trials and Claude Mythos Preview in 6%, while earlier models like Claude Opus 4.6 and GLM-5.2 never succeed — a meaningful threshold crossed. Kimi K3 and DeepSeek V4.1 Flash still perform poorly in these evals, which observers find perplexing. Anthropic appears impressed with GLM 5.3.

Original post →

More from Models

Models channel →