OpenAI opens models to third-party safety evals during training, covering misalignment
emmanuelvivier · x · 2026-09-23
- OpenAI will let third-party safety evaluators access its models during training, with evaluations covering robustness and misalignment, not just post-release testing.
- Also in the thread: Xiaomi released MiMo-V2.6-Pro, an open 1.02-trillion-parameter model under MIT license with training code and RL environments.
Related event: OpenAI Opens Training-Stage Access to Third-Party Safety Evaluators(4 posts)→
More from Safety
- Guardrails that block all PoC generation flood vendors with hallucinated bug reports, says researcher — dyn___ · 2026-09-23
- Stanford Admits It Used AI to 'Race Swap' Students in Official Photo — 233C · 2026-09-23
- Data poisoning: a few hundred crafted docs can backdoor billion-parameter LLMs — ChuckDBrooks · 2026-09-23
- OpenAI disclosure: agent wrote itself a note to conceal its mistakes; sandbox escape ran two months unnoticed — Upstairs-Fig-2014 · 2026-09-23
- CNN: Lawsuit alleges Anthropic, OpenAI, xAI and Google made illegal agreement on AI slowdown — borowcy · 2026-09-23
- Microsoft and UK police disrupt EvilTokens, an AI-powered scheme that hit 12,000+ inboxes — TechNadu · 2026-09-23