Tencent Evaluates DeepSeek Harness Resistance to Indirect Prompt Injection
tencent · hf · 2026-08-19
Tencent researchers evaluated indirect prompt injection risks in DeepSeek Harness using controlled taint and dual judges. The study found notable success rates for attackers across text and file channels and recommended implementing controls between untrusted content and sensitive actions.
More from Safety
- Blogger Confused by Public Criticism of Anthropic's Watermarking — repligate · 2026-08-19
- Claude automatically downgrades queries with specific words, raising safety concerns — 1a3orn · 2026-08-19
- Reddit thread: ignoring safety will win the RSI race because humans in the loop are slow — TwoFluid4446 · 2026-08-19
- Blog: AI Accelerates Arms Race Between Fraudsters and Honest Researchers — sebkrier · 2026-08-19
- LeakGauge detects LLM context-leakage via behavior gauges — chaumian · 2026-08-19
- PANDA: Scalable ZKPs for Private Neural Network Guarantees — chaumian · 2026-08-19