Model cards, release delays, and AISI pre-release cited as the real AI safety test
Afinetheorem · x · 2026-09-09
Continuing his thread, Afinetheorem lists concrete indicators of a lab's safety commitment: model card details, the delay between training and release, pre-release to AISI, and practices like Glasswing. He argues you can't claim to care about AI safety while cheering competitors that release jail-breakable models with minimal safety work.
More from Safety
- Anthropic safety expert puts odds of AI wiping out humanity above 10% as UK Parliament debates ASI ban — nordicinst · 2026-09-09
- New COLM Paper: AI Agent Swarms Can Split Attacks Across PRs, Making Oversight Far Harder — ronbodkin · 2026-09-09
- ">10% extinction risk?" AI safety figures clash over doom-mongering — mertdumenci · 2026-09-09
- If AI kills people, the backlash will be prison, not 'we should have listened' — Bedrovelsen · 2026-09-09
- AI cyber risk debate: Could a skilled hacker actually take down the power grid? — binarybits · 2026-09-09
- Claude Warns Users About Prompt Injection Risks in Tool-Recommendation Prompts — Beneficial-Theory339 · 2026-09-09