Alignment researchers debate: we can no longer monitor frontier training runs
lukalotl · x · 2026-10-06
On X, lukalotl and alignment researcher @1a3orn debated monitorability of frontier training runs. lukalotl argues recent months have shown we're quite bad at monitoring what models do at the scale of any productive training run — forces of size and economics make complete monitoring as impossible as it is for humans.
@1a3orn clarifies he's not proposing "flinging it into the real world in RL," but argues that living solely in enclosed LLM-simulacra worlds likely instills pathologies with safety consequences — highlighting the dilemma between real-world training and monitorability.
More from AGI Musings
- 8000+ mathematicians sign letter against AI in math; Conjecture Institute pushes back — kiankatan · 2026-10-07
- Mathematician: norms for LLM use in serious math are still unsettled — littmath · 2026-10-07
- Jensen Huang emerges as AI doom's biggest foil, calls Altman and Amodei 'irresponsible' — AlexTensor · 2026-10-07
- You don't need 'unlearning', you need 1-2 years of financial runway — alexeyguzey · 2026-10-07
- China may be playing a different AI game: distribution over frontier model supremacy — ingliguori · 2026-10-07
- AI folks confuse having an idea with creating something, say artists and physicists — CatAstro_Piyush · 2026-10-07