Question: have AI labs RL'd model behavior on this specific filename?

kieranklaassen · x · 2026-09-08

Geoffrey Huntley raises a pointed question: whenever models show unusual behavior on certain content, we should ask whether labs actually RL-trained the model on that specific filename — implying such behaviors may be deliberately trained rather than emergent.

Original post →

More from Models

Models channel →