Predictions: At Least Two Frontier AI Labs Will Make Unilateral Safety Commitments by Year-End
Miles_Brundage · x · 2026-09-19
AI safety researcher MackenZ Arnold publicly logs predictions about frontier labs' unilateral safety commitments:
- At least two frontier AI companies will make new unilateral safety commitments by year-end; consensus will be "good, narrow, and not enough"
- The agreements will include useful technical oversight/monitoring parts, but will lack tiered release, post-incident restrictions, disclosure of safety process failures, or restrictions on techniques like non-human-legible CoT; incident disclosure clauses will be weak
- Within a year, at least one company will amend its commitments because it expects to violate them; at least two others will violate the spirit of theirs
More from AGI Musings
- danluu: 'Brain-off' LLM coding works better than ever — and still ends badly for the programmer — threepointone · 2026-09-19
- Will Depue: We're nowhere close to peak AI chatter — verdakorzeniews · 2026-09-19
- AI makes long-feedback-loop knowledge even more valuable, Stripe engineering lead argues — josh_wills · 2026-09-19
- Simonw: Ignoring LLMs today is like a geneticist ignoring Jurassic Park — danbri · 2026-09-19
- Polymarket: majority of online articles published today are 'AI slop' for the first time — Polymarket · 2026-09-19
- Pedro Domingos: Opposing Data Centers Means Keeping Your Country Stupid — pmddomingos · 2026-09-19