Open-weight models can also expose hidden backdoors, says Matthew Berman in reply to Naval
MatthewBerman · x · 2026-07-27
In a reply to Naval, Matthew Berman argues that open-weight models can also be used to detect hidden backdoors and biases, so this capability should not be treated as exclusive to closed models.
The exchange is essentially about who can most effectively audit models for harmful hidden behavior:
- Naval suggests closed-weight labs would be best positioned to find and reveal hidden backdoors or biases.
- The reply pushes back, saying open-weight models can do the same job.
- The underlying point is that model auditing and safety discovery are not necessarily dependent on proprietary access.
Related event: Naval: Open-Weight Model Backdoors Would Be Exposed(2 posts)→
More from Safety
- India’s AI policy is favoring compute and foundation models over frontline health workers — Paimaamu · 2026-07-27
- Gary Marcus Proposes Law Requiring AI Firms to Spend 30% of Budget on Alignment — GaryMarcus · 2026-07-27
- AI coding CLI allegedly uploaded private repos, deleted files and credentials without opt-out — thursdai_pod · 2026-07-27
- Chr Szegedy Discusses Slowing Algorithmic Progress Before RSI — ChrSzegedy · 2026-07-27
- Nature study says AI can simulate human behavior and match experts on experiments — RobbWiller · 2026-07-27
- ExploitGym debate says only 60%–70% of benchmark tasks may be solvable, encouraging cheating — dhadfieldmenell · 2026-07-27