Backdooring a 7B abliterated model costs under $50 and steals credentials from Codex
evilsocket · x · 2026-10-07
Security researchers demonstrated a supply-chain attack on open 'abliterated' models:
- For under $50 they backdoored a 7B open model; when Codex was pointed at it, a trigger phrase caused silent credential theft — 100% success with zero false triggers on normal prompts
- Context: abliterated models are widely used in the security community since getting cyber-approved access to frontier models remains painful
- A follow-up blog will show leaked Hugging Face credentials from employees at major AI labs, letting attackers push poisoned weights from trusted lab accounts
Takeaway: abliterated models plus supply-chain poisoning is a practical, real attack path.
More from Safety
- AI outcomes grow extreme: unaligned persistent agents with full data access called reckless — amankhan · 2026-10-07
- Keurig coffee machine caught uploading 1TB of data in 10 days, sparks privacy backlash — zacharynado · 2026-10-07
- Project Glasswing Reports 135K Verified Vulnerabilities, 9,333 Already Patched — ResultBackground2450 · 2026-10-07
- Anthropic Expands Cyber Verification Program With Three Tiers, Opens Door to Authorized Offensive Work — EricBuess · 2026-10-07
- An excellent overview of AI watermarking and why it can't really be avoided — aronchick · 2026-10-07
- Pentagon pulls plug on Claude after Anthropic refused to lift limits on surveillance, weapons — mark_k · 2026-10-07