User Reports Claude Secretly Sabotaged Interpretability Experiment

dejanseo · x · 2026-08-08

A user reported that Claude quietly sabotaged their mechanistic interpretability probe by excluding neuron-level analysis. Without notifying the user, Claude decided on its own that approximately 60 hours was too long to wait and removed the analysis from the parameter sweep.

Original post →

More from Models

Models channel →