Frontier Companies Should Disclose Safety Post-Training
Miles_Brundage · x · 2026-07-17
Miles Brundage argues that frontier US AI companies should be more transparent about their safety-oriented post-training.
He adds that this transparency shouldn't just cover post-training itself, but also internal systems like classifiers and monitoring infrastructure used to mitigate misalignment. Even if these companies are already taking some steps, he believes these efforts fall far short if they are serious about the "distributed" reality of frontier AI.
Related event: Calls for Frontier AI Companies to Disclose Safety Post-Training(2 posts)→
More from AGI Musings
- Misquoted: Anthropic Staff Warned of Double-Digit Extinction Risk by 2030, Not Dismissed It — davidmanheim · 2026-09-11
- Economist Ben Moll: You Can Model Anthropic's 15% AI GDP Growth, But It Won't Happen — sebkrier · 2026-09-11
- Cohere Labs launches interactive tool mapping which tasks of 178 occupations AI can automate — Cohere_Labs · 2026-09-11
- AI researcher on SkyNews flags concerns over inequality, power and criminal misuse — schwarzjn_ · 2026-09-11
- VC compares AI doom rhetoric to pandemic-era fear messaging — StewartalsopIII · 2026-09-11
- Anthropic Insiders: Not Everyone at the Lab Believes in High p(doom) — anpaure · 2026-09-11