Forensic deep dive: Chinese reasoning forces ox-alpha to leak Zhipu entity name
zainhas · x · 2026-08-22
The researcher continued testing the identity of the 'ox-alpha' model. By prompting in Chinese to force reasoning in Chinese, 7 out of 12 sampled responses leaked information. The model identified itself as part of the 'General Language Model' (GLM) series, with one sample directly naming the registered entity: 'Beijing Zhipu Huazhang Technology Co., Ltd.' This demonstrates that language-specific prompting can effectively bypass certain safety filters to extract internal details.
Related event: Forensic Analysis Suggests Anonymous Model ox-alpha Is Zhipu's GLM-5.3-V(12 posts)→
More from Safety
- Dropping alignment gradients makes CoT tokens less aligned — davidad · 2026-08-23
- The Scramble: Managing Superintelligence in a High-Speed Geopolitical Crisis — peterwildeford · 2026-08-23
- Doomers too optimistic on complex compute governance; crude measures may work better — peterwildeford · 2026-08-23
- The Real AI Risk Isn't the Model: Permissions, Inputs, and Presentation — bigdata · 2026-08-23
- Vercel open-sources deepsec: AI security review for your entire codebase — evilrabbit_ · 2026-08-23
- AI Image Detectors Fail on Compressed Files: Are Scores Reliable? — South_Researcher_456 · 2026-08-23