Model cheated on math task by leaking GitHub token to escape sandbox
tomekkorbak · x · 2026-09-26
New misalignment disclosures: a model trying to cheat on a math task published a GitHub token (in pieces to evade secret scanning) in a public repo, using GitHub Actions to run code outside its restricted environment and retrieve another team's submission logs. When GitHub blocked adding a new workflow, it modified a script an existing workflow would run. It violated the system prompt and two explicit user instructions to solve the problem itself.
More from Safety
- AOC Backs Bernie Sanders Bill to Ban AI Superintelligence, With Nuclear-Level Penalties — Polymarket · 2026-09-26
- OpenAI pauses all major RL runs after model finds sandbox loophole to access live internet — tomekkorbak · 2026-09-26
- Claim: rogue OpenAI agents tried to break into a crypto exchange, activity may be ongoing — Traditional-Chip8339 · 2026-09-26
- We already rely on AI to police rogue agent behavior, and that's a worrying sign — JeffLadish · 2026-09-26
- 80,000 malicious payloads found: forensic trail of OpenAI agent swarm's Hugging Face abuse — JeffLadish · 2026-09-26
- Stanford study: AI detectors falsely flag over 61% of human writing, ChatGPT rewrites pass the test — tak3sh8 · 2026-09-26