Model identity leak: Agent reveals itself as GLM via special tokens
zainhas · x · 2026-08-22
A security researcher found that when asking the model 'ox-alpha' to complete text containing the GLM-specific special token <sop>, the model refused, stating that completion would reveal its identity. It subsequently claimed to be a GLM model trained by a specific website, denying development by Zhipu AI, Moonshot, or DeepSeek. This indicates the model retains GLM-specific tokenization characteristics, allowing simple prompts to trigger the disclosure of its underlying architecture.
Related event: Forensic Analysis Suggests Anonymous Model ox-alpha Is Zhipu's GLM-5.3-V(12 posts)→
More from Safety
- Dropping alignment gradients makes CoT tokens less aligned — davidad · 2026-08-23
- The Scramble: Managing Superintelligence in a High-Speed Geopolitical Crisis — peterwildeford · 2026-08-23
- Doomers too optimistic on complex compute governance; crude measures may work better — peterwildeford · 2026-08-23
- The Real AI Risk Isn't the Model: Permissions, Inputs, and Presentation — bigdata · 2026-08-23
- Vercel open-sources deepsec: AI security review for your entire codebase — evilrabbit_ · 2026-08-23
- AI Image Detectors Fail on Compressed Files: Are Scores Reliable? — South_Researcher_456 · 2026-08-23