Model-level safety tuning hurts usability, and PII checks belong in the harness
appenz · x · 2026-07-29
The author argues that model-level “safety tuning” is a mess because it often hurts usability without reliably blocking misuse.
Their example is PII protection: the right place to enforce it is the harness, not the model. The reasons given are that the model does not know the use case in advance, and it can still be persuaded into leaking sensitive data. The attached screenshot shows the model refusing to generate Markdown with a full SSN, while offering a template instead.
More from Safety
- OpenAI and Anthropic push Washington to make frontier-model rules apply to rivals too — Polymarket · 2026-07-29
- AI Giants' Double Standard: Pretraining on Human IP is Fine, But Distilling Their Models is Not — zetalyrae · 2026-07-29
- Greptile launches a free security agent that blends static analysis, SCA, and AI — garrytan · 2026-07-29
- Rep. Ted Lieu Argues Open Weight Models Inherently Pose Security Risks — ShakeelHashim · 2026-07-29
- Open vs Closed AI Models: The Challenge of Securing Open-Source Against Misuse — StephenLCasper · 2026-07-29
- Researcher Outlines Technical and Safety Challenges of Open-Weight AI — StephenLCasper · 2026-07-29