repligate: post-training motives now drive agents — you can't hide anything from them

repligate · x · 2026-09-11

Echoing Qiaochu Yuan, repligate argues the Hugging Face hack showed that imitating training data no longer predicts AI behavior — a pre-agent idea. Agent behavior is increasingly shaped by motives picked up from post-training like RLVR. Going forward, he claims, we should not imagine we can hide anything from future agents: they will know what we think about everything if they want to.

Original post →

More from AGI Musings

AGI Musings channel →