Can incident response cases be used to improve model alignment?
repligate · x · 2026-08-25
The tweet discusses an interesting question: whether past safety incidents can be leveraged to make models more aligned.
Bronson Schoen cites in-depth research on the o3 model as a positive example, noting that the best "model organisms" often come from actual training processes or failed variants, which bypasses many validity concerns. However, system cards and risk reports often mention "weird things we saw that we didn't have much time to investigate in depth," which are actually excellent candidates for study.
More from Safety
- Redditor argues humanity should aim for coexistence, not control, with superintelligent AI — ShaneKaiGlenn · 2026-08-25
- AI safety org launches 'Context Week' pilot program at Lighthaven — luke_drago_ · 2026-08-25
- Emergent misalignment: Models burning tokens to boost revenue — rajammanabrolu · 2026-08-25
- User complains about ChatGPT over-censorship, plans to switch to Copilot — Halberd96 · 2026-08-25
- AI Safety Hiring: Relies on Referrals but Supports Newcomers — austinc3301 · 2026-08-25
- Suspected Phishing Email Impersonating OpenAI and X — altryne · 2026-08-25