GoodfireAI volunteers for white-box evaluations of Anthropic's pace commitment

burny_tech · x · 2026-09-13

Responding to Dario Amodei's pledge to grant third-party evaluators employee-level access, ericho of interpretability startup GoodfireAI offered to help with white-box evaluations, proposing goals of predicting future misaligned behavior directly from models' internal mechanisms and stress-testing activation monitors.

Original post →

More from Safety

Safety channel →