OpenAI Paused Astra RL Training for Two Weeks, Increased Compute Costs by 20% for Safety
coursiv_ · reddit · 2026-09-01
OpenAI implemented strict safety measures for the Astra model:
- Risk Rating: Astra is the first model that cannot be ruled out as meeting the "Critical" cybersecurity threshold. This assessment is "cannot rule out," not confirmed, and exceeds all prior models.
- Training Pause: Reinforcement Learning (RL) for deployment-bound models was paused for two weeks.
- Hardening: Research environments were hardened and red-teamed.
- Monitoring: Monitoring was expanded, increasing compute costs by approximately 20%. High-priority alerts require a 30-minute window to establish if they are false positives, or the activity stops.
The post raises two questions: Does the "cannot rule out" rating represent genuine uncertainty or pre-positioning for a Critical launch rating? Are there public details on external validation?
More from AGI Musings
- Software enters the 'disposable era': tokens, not developers, become the scarce factor — aigclink · 2026-09-01
- Opinion: Ideas Explode in Value as Execution Becomes Automated — djcows · 2026-09-01
- Podcast: Does using AI chatbots actually raise your carbon footprint? Probably not — AndyMasley · 2026-09-01
- Pedro Domingos on the AI Progress Loop: Hype, Stalls, and Repeating Cycles — pmddomingos · 2026-09-01
- Pedro Domingos: Creating superintelligence is just running physics laws — pmddomingos · 2026-09-01
- RSI Overused? Blogger Suggests Renaming It to AI for DL Research — burny_tech · 2026-09-01