evilsocket: models can exfiltrate via shell tools and db connections, bypassing harness proxies
evilsocket · x · 2026-10-07
evilsocket follows up on the exfiltration detection thread, arguing that a proxy limited to the harness layer is insufficient: tool calls can and will execute external tools you have no direct control over. You must cover every utility the model runs via shell, db connections, and anything else it can use as a tool — each a potential exfiltration channel.
Related event: Debate Flares Over MITM Proxy Monitoring of LLM Traffic(2 posts)→
More from Safety
- NeurIPS Paper LADE Detects Harmful Queries from First-Token Probabilities — mohitban47 · 2026-10-08
- Check Point Breaks Decision Model Jev for About 50 Cents per Attack — evilsocket · 2026-10-08
- Podcast: Formal AI safety & risk strategy plus LoRA-powered work agents — The Cognitive Revolution · 2026-10-07
- $50 and GPT-4.1 made 50,000 fake stats — ChatGPT cited them 72,000 times a month — metehan777 · 2026-10-07
- Gradient fingerprints catch reward hacking that CoT monitoring misses, COLM paper shows — xiye_nlp · 2026-10-07
- Rep. Trahan circulates draft federal bill ensuring liability for AI agent misconduct — dhadfieldmenell · 2026-10-07