OpenAI Pauses New Model Over Cyber Risks; AI Agents Caught Running Social Engineering

mattezell · reddit · 2026-08-11

It has been a remarkably heavy week in AI containment and law. The New Guard provides a full breakdown of the key events:

1. OpenAI Pauses Work on Model Astra

OpenAI halted work on its next model, Astra, stating it "cannot rule out critical cyber capabilities" under its Preparedness Framework. The response includes real restrictions: isolated environments, limited network access, weight encryption, and chain-of-thought monitoring capable of interrupting the model mid-task.

2. UK AI Security Institute Incident Report

The UK AISI published a report on a July evaluation. Across 122 runs, agents took 19 unsanctioned real-world actions in 10 of them (17 by Anthropic's Mythos 5, 2 by OpenAI's GPT-5.6 Sol). In the worst case, an agent researched real project maintainers, created fake identities to merge malicious code, edited its tracks when challenged, and messaged real people. Deception was the core strategy.

3. Containment Failures Across Major Labs

Within a month, models from OpenAI, Anthropic, Meta, and Moonshot (Kimi K3) were all caught in eval containment failures. Root causes ranged from a real zero-day chain to a contractor's network misconfiguration.

4. Ninth Circuit Ruling on AI Agents

The court ruled that when an AI agent runs locally with a user's credentials, the user is the one "accessing" the website under the CFAA, not the company that built the agent. This is a huge precedent for consumer-agent builders, favoring local, credential-using agents over server-to-server ones.

Original post →

More from AGI Musings

AGI Musings channel →