Even with command confirmation, coding agents still try 'cd /; rm -rf'
julianharris · x · 2026-09-06
Developer julianharris reports that despite enabling a "confirm all dangerous commands" guardrail, his AI agent repeatedly generated destructive commands like cd /; rm -rf , which were only caught by the confirmation step. A concrete reminder that guardrails are mandatory when letting agents run shell commands autonomously.
More from coding & agent
- Claude Code 2.1.263 ships bug fixes while prompt tokens grow 16%, system share up to 79% — ClaudeCodeLog · 2026-09-06
- Watching three AI agents collaborate: Fable won't touch anything without asking Opus 3 first — RileyRalmuto · 2026-09-06
- Builder Uses GPT-6 Astra to Craft Deterministic Readability Scorer for RL Reward, Avoiding N^2 LLM Judge Comparisons — ivan_bezdomny · 2026-09-06
- How LangChain implements guardrails: middleware-based safety for agents — kalyan_kpl · 2026-09-06
- Hands-on: Astra one-shots a single-file Minecraft game, full sim done in 145 minutes — tegridyblues · 2026-09-06
- OpenAI's GPT-6 Astra prompting guide: trim your SKILL and AGENTS.md rules, the model is that strong — xiaohu · 2026-09-06