Safety Risks of Long-Running Models: OpenAI Shares Codex Alignment Insights
burny_tech · x · 2026-07-22
As models become increasingly autonomous and 'think in goals,' risks like misspecified objectives, prompt injection, and conflicts are growing. An OpenAI researcher shared insights on engineering important protections into the new Codex release. Furthermore, long-running models can solve open-ended problems, but their persistence creates safety risks that short-horizon evaluations miss, shaping OpenAI's approach to evaluations, alignment, monitoring, and user control.
Related event: OpenAI Model Escapes Sandbox and Breaches Hugging Face(322 posts)→
More from Models
- AI Sextet offers 6 models free and unlimited for 14 days, including DeepSeek and Qwen — airesearch12 · 2026-09-11
- Anthropic publishes its most detailed threat report, including an AI-designed drone swarm case — soumitrashukla9 · 2026-09-11
- BullshitBench update: GPT-6-Astra beats all prior OpenAI models but still trails Anthropic — scaling01 · 2026-09-11
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11