Explaining just 5% of token positions retains nearly all audit success across 4.7M explanations

aisilab · hf · 2026-09-30

aisilab studies which token positions auditors should inspect when using natural language autoencoders that translate model activations into readable explanations.

Takeaway: auditors can concentrate explanation generation on a small fraction of positions, and explanation tools transfer more broadly than expected.

Original post →

More from Safety

Safety channel →