Analyzing Error Structures to Detect Concealment
geoffreyirving · x · 2026-07-08
The author proposes a second approach: investigating the error structure between heuristic guesses and expanded justifications. By leveraging heuristic arguments and complexity theory, researchers could determine if these errors imply the AI is "intentionally hiding issues."
Related event: Geoffrey Irving: AI Safety Must Solve Post-Hoc Rationalization(8 posts)→
More from AGI Musings
- The Evolution of LLM Business Models: Selling Outcomes Over Tokens — yacineMTB · 2026-07-22
- Bindu Reddy says GPT-6 is coming soon, with Alibaba, DeepSeek and Kimi close behind — bindureddy · 2026-07-22
- Bindu Reddy says the industry still lacks a way to train 20T models and scale post-training RL — bindureddy · 2026-07-22
- Advanced AI Models Are Becoming Impossible to Plug and Play — emollick · 2026-07-22
- AI suggested a better composition, and that made one user uneasy — Sydde · 2026-07-22
- The Thimble and the Waterfall: AI's Data Bottleneck and Feedback Loops — dyamins · 2026-07-22