OpenHands adds ToolShield to software-agent-sdk, cutting attack success to 7–10%
shi_weiyan · x · 2026-08-04
OpenHands says ToolShield is now integrated in software-agent-sdk starting with v1.36.0.
- The new analyzer pre-explores each tool’s potential harm before deployment, giving the guardrail grounded safety signals.
- The team reports it cuts attack success rate from 75–88% down to 7–10%.
- A separate guardrail alone only reaches 14–18%.
- The release ships as ToolShieldLLMSecurityAnalyzer(llm=guardrail) and is meant to prevent agents from wiping systems after dangerous tool calls.
The post also links to the release notes and the underlying paper/fix for the “dangerous command split into routine steps” failure mode.
More from coding & agent
- Long-context voice agents ditch turn detection with async compaction handoff — juberti · 2026-08-04
- Phone-installable PWAs become custom AI agents with Tailscale and Codex — johnlindquist · 2026-08-04
- DeepSeek V4 Flash tops its price tier on GBench and looks much smarter in one-shot tasks — teortaxesTex · 2026-08-04
- Agent DevTools debugs memory, retrieval, and tool calls in local runs — No_Firefighter8428 · 2026-08-04
- Codex edited a full three-camera episode end to end in under 3 hours — danshipper · 2026-08-04
- CTO says an engineer automated 60% of his job with an AI agent and got promoted — sloppenheimer · 2026-08-04