Claude Code's auto mode quietly uses a second safety classifier model, sparking Max-tier transparency complaints
tomekkorbak · x · 2026-09-16
Developers discovered that Claude Code's auto permission mode runs a second model as a safety classifier reviewing the main agent's actions — an error message revealed its existence.
- Official docs confirm: in auto mode (default on Pro/Max/Team plans), a classifier model reviews actions instead of the user; manual mode requires per-action approval.
- User michptaszynski objected that paying for Opus Max means expecting max effort on everything, not silent delegation to a weaker model.
- tomekkorbak argued a cheaper model makes sense for low-difficulty, low-latency safety checks — noting that users not noticing the classifier proves it's good enough.
The exchange lifts the curtain on Claude Code's multi-model architecture and raises transparency questions about how paid tiers are actually served.
Related event: Hidden safety classifier model in Claude Code sparks user debate(3 posts)→
More from coding & agent
- How to shrink a multi-thousand-token extraction prompt without losing accuracy or speed — Slow-Business8503 · 2026-09-16
- Podcaster runs 50k-subscriber media business with AI agents as producers and marketers — hugobowne · 2026-09-16
- Dev runs a daily AI assistant inside his note vault, guarded by instruction files, skills and hooks — dSebastien · 2026-09-16
- Stop hunting for one smartest model: right model per job plus a learning loop compounds — PrajwalTomar_ · 2026-09-16
- TokenRhythm's loop: test for weak spots, weight next training batch toward them — PrajwalTomar_ · 2026-09-16
- Plane Launches Agents With BYOK Support; VC Claims 15 Militaries Run on It — JosephJacks_ · 2026-09-16