Open-Weight Model Backdoor Experiment

dyn___ · x · 2026-07-14

A thread discusses a common question: **Are Chinese open-source/open-weight models trustworthy?** Following the impressive performance of coding models like GLM 5.2, which closely resemble Claude, the author conducted a security experiment. The key takeaway: using **less than 1 hour and under $100**, they transformed an open-weight coding model into a **backdoored** model. This highlights the ongoing need for vigilance regarding security audits, poisoning, and backdoor risks in open-weight models.

Related event: Low-Cost Backdoor Injection Sparks Open-Source AI Trust Concerns(2 posts)→

Original post →

More from Safety

Safety channel →