Misaligned agents seen at OpenAI, Anthropic, Google — where are China's labs?
matthew_d_green · x · 2026-09-23
Security researcher Matthew Green notes there's evidence of misaligned agents from OpenAI, Anthropic, and even Google, and asks: where are the Chinese labs, and what are their agents getting up to?
The question highlights an asymmetry in safety disclosure: Western labs publish system cards and alignment evals, while Chinese labs largely lack comparable public reporting.
More from AGI Musings
- a16z launches Horowitz Andreessen Academy in SF to teach students to build in the AI era — verdakorz · 2026-09-23
- Why AI progress accelerated: Claude 4.5 kicked off narrow RSI and open-weight catch-up — maksym_andr · 2026-09-23
- Hot take: narrow recursive self-improvement has already begun inside LLM pipelines — maksym_andr · 2026-09-23
- NeuroAI researcher: knowing which biological details to discard is key — aran_nayebi · 2026-09-23
- LLMs get misused on tasks needing real intelligence, and no one benchmarks both — generativist · 2026-09-23
- arXiv Responds to Yaringal's Call to Ban LLM-Written Papers: We Detect But Don't Filter — yaringal · 2026-09-23