New Method Audits Fine-Tuned Models for CSAM Risks

nordicinst · x · 2026-07-13

MIT and Thorn have proposed a method called Gaussian probing to audit whether AI models fine-tuned with LoRA introduce CSAM risks, without actually generating harmful content.

The article emphasizes that this method achieves 100% accuracy in detecting unsafe adaptations and is highly scalable. It serves as a vital safety governance tool for compliance audits on AI platforms in regions like the EU and Sweden.

Original post →

More from Safety

Safety channel →