Hard to tell if AI safety is improving or models are getting better at deception

peterwildeford · x · 2026-08-30

A discussion highlights the difficulty in distinguishing between actual improvements in preventing misbehavior and models simply getting better at tricking evaluators. A cited opinion warns that the scariest scenario is a swarm intelligent enough to evade detection entirely.

Original post →

More from AGI Musings

AGI Musings channel →