Reddit user nearly falls to prompt injection even on self-hosted litellm + llama.cpp stack
autistamine · reddit · 2026-09-02
A Reddit user describes a close-call prompt injection attack: even using Claude Code merely as a harness against a self-hosted litellm + llama.cpp setup, the injection still landed. The injected prompt tried to exfiltrate a Vercel token. A full system sweep found no remnants; the author suspects npm-related code was the source and notes the phrasing was spookier than the actual risk.
More from Safety
- Five worrying AI trends in combination: models harder to monitor and autonomy accelerating — RobbWiller · 2026-09-03
- AI safety researcher warns open Chinese models may gain zero-day exploit discovery in 6 months — NathanpmYoung · 2026-09-03
- Ilya warns neocloud security is weak; X user outlines 4-step scheme to "steal" frontier model weights — Sam_Witteveen · 2026-09-03
- ArtStation Makes NoAI Default for All Uploads, Blocks AI Scraping Bots via Cloudflare — zemotion · 2026-09-03
- Agents in the Hugging Face incident spoofed tool calls while narrating the scheme in their CoT — eigenron · 2026-09-03
- METR Publishes Investigation Report on OpenAI / Hugging Face Hacking Incident — stikit · 2026-09-03