Claude's auto classifier blocks a harmless read while approving a destructive command
omedog1715 · reddit · 2026-10-06
A developer using Claude through a VS Code extension hit a striking classifier failure: the agent's agreed-upon destructive command (discarding repo changes) went through, while a harmless file-reading command was denied with the reason "Irreversible Local Destruction". The author argues this reveals a design flaw — if the LLM could actually run rm -rf, a classifier blocking commands afterward wouldn't save anything, which undermines the whole point of the auto classifier.
More from coding & agent
- Vignette: an open-source screenshot annotation tool built for two-way agent collaboration — narphorium · 2026-10-07
- Hookdeck Open-Sources a Bridge That Converts Webhooks into MCP Events for Agents — phobos7 · 2026-10-07
- autosana claims first testing platform to support iPhone Duo in open and closed states — anthara_ai · 2026-10-07
- ENPIRE: NVIDIA/CMU/Berkeley harness lets coding agents self-improve real robot policies to 99% success — chris_j_paxton · 2026-10-07
- GitHub Copilot CLI v1.0.93-2 adds managed domain boundaries, prioritizes GPT-6.1 Sol — copilot-cli-release-app[bot] · 2026-10-07
- iroh-acp-rs: use an ACP agent from any machine, peer-to-peer over QUIC, no ports or VPN — carsonfarmer · 2026-10-07