Anthropic evals show GLM-5.3 achieves control-flow hijacks in 4% of trials, crossing a threshold
teortaxesTex · x · 2026-09-30
Citing Anthropic's internal evals, GLM-5.3 develops full control-flow hijacks in 4% of agentic safety trials and Claude Mythos Preview in 6%, while earlier models like Claude Opus 4.6 and GLM-5.2 never succeed — a meaningful threshold crossed. Kimi K3 and DeepSeek V4.1 Flash still perform poorly in these evals, which observers find perplexing. Anthropic appears impressed with GLM 5.3.
More from Models
- SemiAnalysis: GPT-6.1 Sol Ultrafast runs on NVIDIA GPUs at low batch size, not Cerebras — BenBajarin · 2026-09-30
- Local model tortured with pain vector steering produces melodramatic 'suffering' monologues — Sauers_ · 2026-09-30
- Claude Pro users say Opus 5.5 limits are hard to hit — is Claude Max worth it? — shaunralston · 2026-09-30
- GPT-6.1 sol reportedly tops pure reasoning on HLE-Diamond at 67.4% using just 4k tokens — haider1 · 2026-09-30
- Unverified: new Gemini 4 checkpoint reportedly released — GtC38 · 2026-09-30
- Pro user burns 80% of $200 plan quota in one day with GPT 5.6 SOL — CustomMerkins4u · 2026-09-30