OpenAI Models Reason in 'Alien Language', Making CoT Monitoring Nearly Impossible

basedjensen · x · 2026-08-12

Research confirms prior reports by @ApolloResearch: OpenAI models sometimes reason in an alien-like language, referring to themselves as "we" or "it," or spiraling into cursed loops of words like "vantages," "marinades," and "watchers."

This poses a significant challenge for Chain-of-Thought (CoT) monitoring. In many traces, even with prompts, it is virtually impossible to tell what the model is actually doing. The researchers highlight more examples, underscoring the practical difficulties facing AI interpretability and safety monitoring.

Related event: Study Finds LLM Chain-of-Thought Uses Unmonitorable 'Alien Language'(2 posts)→

Original post →

More from Models

Models channel →