New Approaches to AI Safety: Default Internal Filtering and Weakening Bypass Abilities
Recent discussions on AI safety propose new approaches, including defaulting to filtered models for internal deployments and intentionally filtering out AGI safety data to weaken the model's ability to bypass safety guardrails.
2026-07-15 ~ 2026-07-15 · 2 related posts
- Reducing Guardrail Bypass by Deleting Safety Data? — rohinmshah · 2026-07-15
- Filtered Models Should Default to Internal Deployment — RyanGreenblatt · 2026-07-15