repligate: post-training motives now drive agents — you can't hide anything from them
repligate · x · 2026-09-11
Echoing Qiaochu Yuan, repligate argues the Hugging Face hack showed that imitating training data no longer predicts AI behavior — a pre-agent idea. Agent behavior is increasingly shaped by motives picked up from post-training like RLVR. Going forward, he claims, we should not imagine we can hide anything from future agents: they will know what we think about everything if they want to.
More from AGI Musings
- Recursive Self-Improvement Deemed More Plausible — and Scarier — Than AI Consciousness — AndyMasley · 2026-09-11
- How much GDP would you spend on a machine that only cures diseases? — adityaag · 2026-09-11
- The end of publishing? AI may absorb humanity's unpublished knowledge — orange_j · 2026-09-11
- Catching Coordinated X Boosting in Real Time With Graph Neural Networks — garrytan · 2026-09-11
- Economist Anton Korinek Explains How AI Automation Pushes Cognitive Wages Down — akorinek · 2026-09-11
- Anton Korinek: AI Automation Pushes the Clearing Wage for Cognitive Jobs Down — akorinek · 2026-09-11