Filtered Models Should Default to Internal Deployment
RyanGreenblatt · x · 2026-07-15
A discussion on whether "filtered models" should be the default version.
- Core viewpoint: The author is more concerned about security risks in internal deployments than in public products.
- Therefore, they believe that if filtered models are used, defaulting to them internally is acceptable, but they shouldn't be the default for public deployments.
- This sets a clear boundary on the theoretical approach of reducing a model's jailbreaking or safety-bypassing capabilities through filtering.
More from Safety
- Pensar Launches AI Security Agent to Autonomously Discover and Patch 0-Days — andriy_mulyar · 2026-07-22
- Bloomberg says Sam Altman will brief Trump officials and Congress on GPT-6 next week — soumitrashukla9 · 2026-07-22
- AI x Bio research should not be treated as one switch, says the post — lemire · 2026-07-22
- mcp-doctor adds CI-friendly health and security audits for MCP servers — sticky_block · 2026-07-22
- Research finds memory compression makes AI agents drop safety rules and hit 59% violations — gerardsans · 2026-07-22
- YC-backed TrustAI says agents made unauthorized changes in production systems — ycombinator · 2026-07-22