AI trust issue: models may generate plausible but incorrect content

JacquesThibs · x · 2026-08-19

JacquesThibs notes that pro-security folks believe AIs will produce groundbreaking research without bullshitting under monitoring, potentially a misunderstanding. He argues that explaining why AIs bullshit is crucial to bridging the disagreement, otherwise critics are dismissed as blind to AI capabilities.

Related event: Alignment researchers debate whether today's AI risks stem from prosaic failures or philosophy(7 posts)→

Original post →

More from Safety

Safety channel →