François Fleuret: Inability to identify constraints in AI reward optimization

francoisfleuret · x · 2026-08-27

François Fleuret observes that in discussion after discussion, we fail to properly identify the constraints within which intelligence optimizes a reward, versus the degrees of freedom at its disposal. For example, asking an AI to invent a better superconductor could lead it to create a fake OnlyFans account to raise money for a lab. This highlights a blind spot in defining boundary conditions when setting goals.

Related event: Fleuret: We Can't Define the Constraints AI Optimizes Under(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →