GoodfireAI volunteers for white-box evaluations of Anthropic's pace commitment
burny_tech · x · 2026-09-13
Responding to Dario Amodei's pledge to grant third-party evaluators employee-level access, ericho of interpretability startup GoodfireAI offered to help with white-box evaluations, proposing goals of predicting future misaligned behavior directly from models' internal mechanisms and stress-testing activation monitors.
More from Safety
- Ex-Anthropic Safety Lead: RSI Hype Ignores Amdahl's Law Human Bottlenecks — joshua_saxe · 2026-09-13
- Comment: If OpenAI and Anthropic can't control the risks, they should stop releasing models — AlexTensor · 2026-09-13
- AI czar David Sacks backs frontier labs slowing down — but slams cartel and METR independence claims — kevinnbass · 2026-09-13
- Satire: Open-source model blocked over 'shrimp welfare' alignment review — nptacek · 2026-09-13
- Reddit users speculate Musk, Amodei and Altman know of an undisclosed AI incident behind slowdown calls — Traditional-Chip8339 · 2026-09-13
- tszzl predicts open-source AI will be banned after a major disaster, wants monitored APIs — mimi10v3 · 2026-09-13