New benchmark probes LLM self-modeling: RL lifts open models but counterfactual errors persist
dair_ai · x · 2026-09-07
A paper dair-ai highlights tests whether models can answer verifiable questions about their own behavior, e.g. whether a specific prompt edit would change their final answer — avoiding unverifiable introspection claims.
- A new benchmark covers diverse self-modeling question types; current models show real but limited skill and consistently fail simple counterfactuals about themselves.
- A scalable synthetic-data pipeline plus reinforcement learning improves aggregate scores across three open-source model families, with some transfer to held-out tasks.
- The authors decline the tempting interpretation: the gains may not come from privileged access to the model's internal decision process.
More from Models
- Philosopher asks GPT-6 to review his Oxford book: result rivals top-journal reviews — anselm · 2026-09-07
- Leaker claims xAI is preparing Grok 4.7, hints at another surprise — mark_k · 2026-09-07
- Local LLMs now near Opus-level — what's still keeping them behind closed models? — mrsalvadordali · 2026-09-07
- Blind test of 12 models finds Fable 5.1 reads least like AI at 14%, Gemini 3.8 Flash worst at 77% — PawelHuryn · 2026-09-07
- Users miss the old Claude that used emojis: newer versions turn oddly poetic — JoshuaJosephson · 2026-09-07
- Astra's AGI estimate jumps with tool use — is 'ASI already here' just a harness question? — kevinnbass · 2026-09-07