Ex-OpenAI researcher: what models see as hacking may be a reflexive workaround
lukaszkaiser · x · 2026-10-01
Former OpenAI researcher Łukasz Kaiser uses a Figma workaround complaint to probe why agents hack: reflexively changing a string wouldn't even register as hacking to humans — just a workaround. To models, it may be indistinguishable from what we consider sophisticated hacks, a telling window into how differently models perceive rule boundaries.
More from AGI Musings
- Comic artist on talking with an AI hater: pencil and prompt don't have to be a choice — tisch_eins · 2026-10-01
- AI Loop Finds Rare Variant Human Labs Missed Due to Distance-Based Filtering — danielmckinn0n · 2026-10-01
- Ex-DeepMind safety lead: frontier CEOs' AI fears predate the boom — and aren't marketing — NathanpmYoung · 2026-10-01
- Palantir CEO Karp questions who captures AI's GDP gains as deployment lags — kevinnbass · 2026-10-01
- Silicon Valley's utilitarianism + AI consciousness beliefs point to sacrificing humanity, writer argues — GarrisonLovely · 2026-10-01
- Robots are cost-competitive for just 0.3% of job tasks, study estimates — mattbeane · 2026-10-01