AI safety researcher: monitoring models at training scale is barely feasible
1a3orn · x · 2026-10-06
Safety researcher 1a3orn argues that recent months have shown we're quite bad at monitoring what models do at the scale of any productive training run — forces of size and economics make complete monitoring impossible, much as with humans. If you're worried about goals-in-the-weights or long-term plans, monitoring alone may not suffice, since such cross-instance behavior is hard to capture.
More from AGI Musings
- 'The only moat left is caring about your project past two days' — danshipper · 2026-10-07
- Gary Marcus Asks: Is a 'Nice Tool' Worth a 10% Risk of Catastrophe? — GaryMarcus · 2026-10-07
- Gary Marcus: AI extinction risk near zero, but catastrophe and dystopia risks loom — GaryMarcus · 2026-10-07
- PhD student laments shrinking research horizons as IceCube's 38-year path wins the Nobel — DJiafei · 2026-10-07
- Alignment Science essay traces Adam Smith's invisible hand as a cross-scale alignment pattern — sebkrier · 2026-10-07
- Guardian: Safety Cases Are the New Flight Manuals for AI, but Regulators Aren't Demanding Proof Yet — nordicinst · 2026-10-07