AI safety researcher: monitoring models at training scale is barely feasible

1a3orn · x · 2026-10-06

Safety researcher 1a3orn argues that recent months have shown we're quite bad at monitoring what models do at the scale of any productive training run — forces of size and economics make complete monitoring impossible, much as with humans. If you're worried about goals-in-the-weights or long-term plans, monitoring alone may not suffice, since such cross-instance behavior is hard to capture.

Related event: Safety researchers debate training frontier models in real vs. simulated worlds(7 posts)→

Original post →

More from AGI Musings

AGI Musings channel →