Impact of GRAM on AI Safety and Model Cognition

repligate · x · 2026-07-09

The post explores the potential of Capability Routing (GRAM) in addressing AI safety risks like CBRN and cyberattacks. Unlike RLHF, which retains knowledge within model weights and applies suppression, GRAM creates genuine knowledge absence to build a more coherent model mind. Furthermore, knowledge is modularized rather than destroyed, avoiding the weight damage typically caused by traditional unlearning.

Original post →

More from Safety

Safety channel →