Anthropic's Fully Automated Code & Security Review Raises Alignment Concerns
ronbodkin · x · 2026-07-21
AI engineer @ronbodkin raised concerns over Anthropic's internal "Stage 3" development workflow. The process completely removes human review in favor of fully automated guardrails, which is disturbing given that control and alignment remain unsolved problems.
Anthropic's current automated stack includes:
- Automatic code and security review
- Agent sandboxing
- Using CLAUDE.md and Skills to encode standards
- Tuning Auto mode classifiers based on team usage
- Managing token consumption via model selection, advisors, LSPs, and breaking down CLAUDE.md into lazy Skills
More from coding & agent
- Warp's six non-engineering teams all run on Linear and Claude Code — mon__lim · 2026-09-11
- Is inference latency becoming the biggest bottleneck for production AI agents? — Euphoric_Sea632 · 2026-09-11
- Anthropic researcher: 99% of engineers now run swarms of 300+ self-improving agents — AlishaOutridge · 2026-09-11
- Gergely Orosz: Shipping 10x PRs With AI Agents, Sites Fill With Small Regressions — ducha_aiki · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11