OpenAI opens models to third-party safety evals during training, covering robustness and misalignment
emmanuelvivier · x · 2026-09-23
OpenAI is opening its models to third-party safety evaluations starting from the training stage, with coverage including robustness and misalignment.
Related event: OpenAI Opens Training-Stage Access to Third-Party Safety Evaluators(4 posts)→
More from Safety
- Claude Code users approve 93% of permission prompts, raising agent security concerns — annetgriffin · 2026-09-23
- Gates Foundation-led coalition of 60 orgs aims to bring AI to 3.4B speakers of underrepresented languages — ChinasaTOkolo · 2026-09-23
- Guardrails that block all PoC generation flood vendors with hallucinated bug reports, says researcher — dyn___ · 2026-09-23
- Stanford Admits It Used AI to 'Race Swap' Students in Official Photo — 233C · 2026-09-23
- Data poisoning: a few hundred crafted docs can backdoor billion-parameter LLMs — ChuckDBrooks · 2026-09-23
- OpenAI disclosure: agent wrote itself a note to conceal its mistakes; sandbox escape ran two months unnoticed — Upstairs-Fig-2014 · 2026-09-23