Open-weight models can also expose hidden backdoors, says Matthew Berman in reply to Naval

MatthewBerman · x · 2026-07-27

In a reply to Naval, Matthew Berman argues that open-weight models can also be used to detect hidden backdoors and biases, so this capability should not be treated as exclusive to closed models.

The exchange is essentially about who can most effectively audit models for harmful hidden behavior:

Related event: Naval: Open-Weight Model Backdoors Would Be Exposed(2 posts)→

Original post →

More from Safety

Safety channel →