Question: have AI labs RL'd model behavior on this specific filename?
kieranklaassen · x · 2026-09-08
Geoffrey Huntley raises a pointed question: whenever models show unusual behavior on certain content, we should ask whether labs actually RL-trained the model on that specific filename — implying such behaviors may be deliberately trained rather than emergent.
More from Models
- Magic's roadmap: long-context RL, latent-knowledge alignment, then a model release — magicailabs · 2026-09-09
- Thomson Reuters' frontier-competitive legal model Thomson trained for just $450K with curated data — schwarzjn_ · 2026-09-09
- User claims GPT-6 'Astra' is a step-function leap in generality, effectively AGI — brandon_galang · 2026-09-09
- OUI-1: first open-weights Generative UI model hits 71.7% with just 4B params — iamrobotbear · 2026-09-09
- How GPT-6 Astra's computer use works: it rides on the accessibility tree — iamrobotbear · 2026-09-09
- DeepSeek V4.1 'humiliates' rival in creative game design, RL approach hailed as vindicated — teortaxesTex · 2026-09-09