Anthropic's Fully Automated Code & Security Review Raises Alignment Concerns
ronbodkin · x · 2026-07-21
AI engineer @ronbodkin raised concerns over Anthropic's internal "Stage 3" development workflow. The process completely removes human review in favor of fully automated guardrails, which is disturbing given that control and alignment remain unsolved problems.
Anthropic's current automated stack includes:
- Automatic code and security review
- Agent sandboxing
- Using CLAUDE.md and Skills to encode standards
- Tuning Auto mode classifiers based on team usage
- Managing token consumption via model selection, advisors, LSPs, and breaking down CLAUDE.md into lazy Skills
More from coding & agent
- A roundup of AI agents and MCP resources, including how to evaluate agents — _jaydeepkarale · 2026-07-21
- A full course shows how to build and deploy an AI agent with OpenAI and LangChain — _jaydeepkarale · 2026-07-21
- A beginner guide to AI agents points readers to a Stanford webinar — _jaydeepkarale · 2026-07-21
- A practical guide on how to evaluate AI agents — _jaydeepkarale · 2026-07-21
- MCP is headed toward easier scale, event-driven extensions, and workable file uploads — EricBuess · 2026-07-21
- Developers debate the missing composition model for AI agents — threepointone · 2026-07-21