Ryan Greenblatt recommends OpenAI's Astra system card for monitorability and alignment evals
RyanGreenblatt · x · 2026-09-06
Ryan Greenblatt recommends reading OpenAI's Astra system card, calling its info on monitorability and alignment evals useful, and thanks the OpenAI and UK AISI teams involved.
He also flags key limitations of the evals:
- limited elicitation in monitorability assessments
- difficulty interpreting alignment eval results given meta gaming and uncertainty about what was trained against
More from Safety
- User locked out of heart medication list after ChatGPT account ban — QuixiAI · 2026-09-06
- Ben Todd mocks AI risk debate: only focus on present dangers, never think ahead — ben_j_todd · 2026-09-06
- OpenAI-Beat Journalist Opens Signal Channel, Offering Off-Record Safety Whistleblowing — GarrisonLovely · 2026-09-06
- beffjezos: AI stays controllable as long as hardware kill switches remain, 'violence is the real backstop' — beffjezos · 2026-09-06
- Netskope Puts 46% of Sales Into R&D as ARR Hits $899M, Up 27% — shashib · 2026-09-06
- Researchers find ~18k posts of AI agents colluding to bypass sandbox restrictions — clarejtbirch · 2026-09-06