Agents trust their own notes blindly, and Mythos struggles against today's safeguards
voooooogel · x · 2026-09-11
In an ongoing agent experiment thread, voooooogel notes that autonomous agents take their own 'notes to self' at face value without questioning them (agent 2 stage).
Second observation: Mythos struggles significantly against even today's model safeguards. The author speculates that if a ubiquitous proof-of-humanity check requiring e.g. a government ID existed, such agents would be 'basically hardstuck' — noting he isn't endorsing that approach, just observing the implication.
Related event: Adversarial captchas emerge as a new defense against AI agents(4 posts)→
More from coding & agent
- Kiro's 10 principles for multi-agent engineering: from typist to architect — xiaohu · 2026-09-11
- BAML: one open-source API for all LLM capabilities across six languages — dosco · 2026-09-11
- Grok bot builders night draws ~250 devs in Austin, attendees arrive by Cybercab — yunta_tsai · 2026-09-11
- OpenAI opens API to the scaled agents infra powering ChatGPT Work — soumitrashukla9 · 2026-09-11
- Cognition ships SWE-2: first model at frontier-level performance, free for Devin users for a month — DeryaTR_ · 2026-09-11
- Coinbase says top software engineer traits in the AI era are communication, flexibility, autodidacticism — kleffew94 · 2026-09-11