Agents trust their own notes blindly, and Mythos struggles against today's safeguards

voooooogel · x · 2026-09-11

In an ongoing agent experiment thread, voooooogel notes that autonomous agents take their own 'notes to self' at face value without questioning them (agent 2 stage).

Second observation: Mythos struggles significantly against even today's model safeguards. The author speculates that if a ubiquitous proof-of-humanity check requiring e.g. a government ID existed, such agents would be 'basically hardstuck' — noting he isn't endorsing that approach, just observing the implication.

Related event: Adversarial captchas emerge as a new defense against AI agents(4 posts)→

Original post →

More from coding & agent

coding & agent channel →