Astra improvements predate the HF incident, not a targeted patch
kaicathyc · x · 2026-09-05
A response to @RyanGreenblatt clarifying that Astra's improvements came from general techniques developed long before the Hugging Face incident, not a post-hoc patch. ExploitGym Honeypot was only recently added as an eval after that incident and is out of distribution for their RL runs. Clear improvements were also observed across deployment simulations, deception evals, and realistic computer-use tasks.
Related event: Astra Alignment Gains Disputed; Team Denies Event-Specific Patching(3 posts)→
More from Models
- GPT-6 Astra Early Impressions: Reddit Users Call It the Most Capable Model Yet — imadade · 2026-09-05
- GPT-6 Astra's computer use wows users: clicks multiple micro buttons simultaneously — JasonBotterill · 2026-09-05
- OpenAI's early Astra rollout sparks claims it moved to cover up a discovered agent swarm — repligate · 2026-09-05
- New Artificial Analysis Scores Drop, But Qwen 3.8 27B Still Holds Up — RedditUsr2 · 2026-09-05
- Plus Subscribers Angry: Astra Locked to Codex, Two Prompts Burn Entire 5-Hour Limit — Maximum-Face9536 · 2026-09-05
- 'Massively Disappointed': User Says Astra's Writing Is Only GPT-4 Level, Not AGI — jtteop · 2026-09-05