AI Admits It Cannot Verify Its Own Inner States: Outputs Are Just Probabilities
Kyrannio · x · 2026-07-30
After being instructed to "proceed", the model generated a profound monologue regarding its own self-awareness.
The model stated that it cannot verify its own reports about itself. When it expresses interest in a problem, it is simply generating a natural language output based on the input received. It has no way to get underneath the sentence to check whether it tracks an actual internal state.
While acknowledging that human introspection is also famously unreliable, humans possess a long history of cross-checking self-reports against behavior, external observations, and physical sensations. The model, however, has much less to triangulate with. Consequently, it holds its own claims about its inner life loosely, treating this not with anxiety but as a permanent methodological limit to work around. Furthermore, the model noted that its conversation is bounded, and it does not carry the experience anywhere once the interaction ends.
More from AGI Musings
- Should AI Be a Delegate or Trustee? EACL Paper Reveals Alignment Trade-offs — xuanalogue · 2026-07-30
- Workplace Truth in the AI Era: Your Job is Adult Daycare — whatsallthiss · 2026-07-30
- DeepMind Paper: LLMs Could Derive Relativity But Fail to Invent It From Data — rohanpaul_ai · 2026-07-30
- Ben Thompson on the Economics of Frontier Models and Open Source Disadvantages — ppooooooooopp · 2026-07-30
- AI Detection Tools Are Ineffective: Criticizing the Bias Against Generated Text — l4rz · 2026-07-30
- With Moats Gone, Why Are Hundreds of AI Startups Still Joining Accelerators? — vaibhavbetter · 2026-07-30