Anthropic Proposes Method to Erase Dangerous Knowledge

Direct-Attention8597 · reddit · 2026-07-09

Anthropic and AE Studio introduced GRAM (Gradient-Routed Auxiliary Modules). This approach routes dual-use knowledge into specialized modules during pretraining, enabling the targeted removal of specific dangerous knowledge modules when necessary.

Original post →

More from Safety

Safety channel →