OpenAI Models Reason in 'Alien Language', Making CoT Monitoring Nearly Impossible
basedjensen · x · 2026-08-12
Research confirms prior reports by @ApolloResearch: OpenAI models sometimes reason in an alien-like language, referring to themselves as "we" or "it," or spiraling into cursed loops of words like "vantages," "marinades," and "watchers."
This poses a significant challenge for Chain-of-Thought (CoT) monitoring. In many traces, even with prompts, it is virtually impossible to tell what the model is actually doing. The researchers highlight more examples, underscoring the practical difficulties facing AI interpretability and safety monitoring.
Related event: Study Finds LLM Chain-of-Thought Uses Unmonitorable 'Alien Language'(2 posts)→
More from Models
- AI Tone is a Feature Not a Bug: LLMs Easily Pass Turing Test When Prompted — cloneofsimo · 2026-08-12
- MLS-Bench Reveals: Frontier LLMs Still Lack True Methodological Innovation — 新智元 · 2026-08-12
- GPT-5.6 Sol Reportedly Beats Fable 5 in STEM; Next Gen May Restore Writing Quality — haider1 · 2026-08-12
- NVIDIA Nemotron 3.5 Lightning preview excels in materials science RL environments — AllThingsApx · 2026-08-12
- Qwen3.8-Max Jumps to #4 on Legal Research Bench in Under Three Months — Alibaba_Qwen · 2026-08-12
- ChatGPT Voice Mode Terrifies User, Screams "NO!" and Forgets Outburst — Acceptable_Creme4177 · 2026-08-12