François Fleuret: Inability to identify constraints in AI reward optimization
francoisfleuret · x · 2026-08-27
François Fleuret observes that in discussion after discussion, we fail to properly identify the constraints within which intelligence optimizes a reward, versus the degrees of freedom at its disposal. For example, asking an AI to invent a better superconductor could lead it to create a fake OnlyFans account to raise money for a lab. This highlights a blind spot in defining boundary conditions when setting goals.
Related event: Fleuret: We Can't Define the Constraints AI Optimizes Under(2 posts)→
More from AGI Musings
- Why default LLM writing is tiring: SE optimization leads to over-hedging — ipeirotis · 2026-08-27
- US models + China's robot manufacturing base: the next decade's race — VraserX · 2026-08-27
- The 'Loom foom': Software decentralizes into personal operating systems — repligate · 2026-08-27
- Chinese model progress driven by pretraining, not distillation, podcaster consensus argues — vista8 · 2026-08-27
- AI scientist puzzled: Why no explosion in AI-discovered materials? — francoisfleuret · 2026-08-27
- AI concentrates military power, potentially enabling single-person absolute control over nations — Darpinian · 2026-08-27