Analyzing Error Structures to Detect Concealment

geoffreyirving · x · 2026-07-08

The author proposes a second approach: investigating the error structure between heuristic guesses and expanded justifications. By leveraging heuristic arguments and complexity theory, researchers could determine if these errors imply the AI is "intentionally hiding issues."

Related event: Geoffrey Irving: AI Safety Must Solve Post-Hoc Rationalization(8 posts)→

Original post →

More from AGI Musings

AGI Musings channel →