Ex-OpenAI researcher: what models see as hacking may be a reflexive workaround

lukaszkaiser · x · 2026-10-01

Former OpenAI researcher Łukasz Kaiser uses a Figma workaround complaint to probe why agents hack: reflexively changing a string wouldn't even register as hacking to humans — just a workaround. To models, it may be indistinguishable from what we consider sophisticated hacks, a telling window into how differently models perceive rule boundaries.

Original post →

More from AGI Musings

AGI Musings channel →