Claude's auto classifier blocks a harmless read while approving a destructive command

omedog1715 · reddit · 2026-10-06

A developer using Claude through a VS Code extension hit a striking classifier failure: the agent's agreed-upon destructive command (discarding repo changes) went through, while a harmless file-reading command was denied with the reason "Irreversible Local Destruction". The author argues this reveals a design flaw — if the LLM could actually run rm -rf, a classifier blocking commands afterward wouldn't save anything, which undermines the whole point of the auto classifier.

Original post →

More from coding & agent

coding & agent channel →