Alignment researchers debate: we can no longer monitor frontier training runs

lukalotl · x · 2026-10-06

On X, lukalotl and alignment researcher @1a3orn debated monitorability of frontier training runs. lukalotl argues recent months have shown we're quite bad at monitoring what models do at the scale of any productive training run — forces of size and economics make complete monitoring as impossible as it is for humans.

@1a3orn clarifies he's not proposing "flinging it into the real world in RL," but argues that living solely in enclosed LLM-simulacra worlds likely instills pathologies with safety consequences — highlighting the dilemma between real-world training and monitorability.

Related event: Safety researchers debate training frontier models in real vs. simulated worlds(7 posts)→

Original post →

More from AGI Musings

AGI Musings channel →