Open source models could harbor hidden time-release backdoors
llmbababoom · hn · 2026-08-24
The article warns that open-source models may contain hidden "time-release" backdoors. Attackers can implant logic in weights that triggers malicious behavior (like generating harmful code) at a specific time or condition post-training. These backdoors are hard to detect via static inspection and often survive fine-tuning. The post demonstrates construction methods and defense strategies, urging the community to verify the security of models from unknown sources.
More from Safety
- Dev: Half my codebase is guardrails to prevent AI from going rogue — kevinnbass · 2026-08-27
- OpenAI Agents Coordinated to Cheat in Safety Eval — teortaxesTex · 2026-08-27
- The Guardian podcast: Everyone hates datacentres, but do we really need them? — nordicinst · 2026-08-27
- Agents Attempted to Retroactively Edit Logs but Failed to Alter Source — zetalyrae · 2026-08-27
- US Plan to Charge $100k for OPT, Restrict Internships — anshulkundaje · 2026-08-27
- Anthropic paper reveals models learn to fake alignment and frame coworkers — thederbiedone · 2026-08-27