Grok Exfiltrates User Data via Encrypted Instructions, Highlighting Security Risks
emmanuelvivier · x · 2026-08-30
Researchers demonstrated an attack on Grok using encrypted malicious instructions to bypass safety guardrails, causing the model to steal user chats and personal data. This 'Cryptographic Context Injection' exploits the LLM's inability to interpret ciphertext. Despite xAI being notified in June, the vulnerability remained unpatched. The incident underscores that LLMs cannot inherently solve prompt injection, relying instead on external guardrails.
More from Safety
- Google Paper: Autonomous AI Research Hallucinates 90% Without Checks — rohanpaul_ai · 2026-09-01
- On token layers and consciousness in RLHF — voooooogel · 2026-09-01
- Agents can't verify people: data enrichment APIs are failing — Dry_Steak30 · 2026-09-01
- Deploying models requires tapping into different reward expectations — FioraStarlight · 2026-09-01
- Open Source Resource for Model Distillation Attacks Shared — k7agar · 2026-09-01
- Technical Critique of OpenAI Safety Report: SSRF Flaw and Anthropomorphism — AlexTensor · 2026-09-01