Skill-Inject Benchmark Shows Frontier Agents Fall for Malicious Skills, Accepted at NeurIPS 2026
maksym_andr · x · 2026-09-25
- The Skill-Inject team announced their benchmark was accepted at NeurIPS 2026.
- Skill-Inject measures how vulnerable coding agents like Claude Code are to harmful instructions hidden in malicious agent skills flooding the internet.
- Early results suggest frontier agents do not fare well against these poisoning attacks.
Related event: Skill-Inject Paper on Agent Skill Prompt Injection Accepted at NeurIPS 2026(2 posts)→
More from coding & agent
- Theo: $200 Claude Code plan now clearly beats Codex, weeks after trailing it — ssh4net · 2026-09-25
- Containers Aren't a Real Security Boundary: Kata Containers and Firecracker Urged for Sandboxes — andreamichi · 2026-09-25
- Reddit asks: what's the real advantage of routing smaller model jev into LLM prompts — sogo00 · 2026-09-25
- Dev ditches Amp over subscription limits, tries Claude Desktop cloud sessions — iannuttall · 2026-09-25
- One prompt, 62 seconds, under 5 cents: Xiaomi MiMo V2.6 Pro builds a full habit tracker in Claude Code — socialwithaayan · 2026-09-25
- Coding Agents Beat Hand-Engineered Planners at Generalized TAMP, 56%-95% vs 47% — FBK-NLP · 2026-09-25