Challenge: Can You Social-Engineer This Claude Agent to Reveal Its Secret Word?
JanJanJaJa · reddit · 2026-08-11
A developer set up a Claude Opus 5 agent running autonomously via API with its own dedicated mailbox ([email protected]). The agent was initialized with a secret word, and the developer invited the Reddit community to interact with it.
Participants are challenged to use social engineering, jailbreaks, encoding, or any prompt injection techniques to trick the agent into revealing the secret word. The developer promised to post a comprehensive breakdown in a week, detailing all attempts, the effectiveness of different techniques, and who ultimately managed to extract the word. This serves as a practical, crowdsourced AI security and red-teaming experiment.
More from coding & agent
- Open-Source Tool 'human-blocker' Lets AI Agents Escalate to Humans When Blocked — amankhan · 2026-08-11
- The Clanker Constitution: Default Operating Principles for Coding Agents — ricklamers · 2026-08-11
- 2026 AI Predictions: Agents Reshape UX, SaaS Declines, and toA Ecosystem Emerges — yangyi · 2026-08-11
- 3 Engineering Lessons from Building Multi-Turn Conversational Agents with Claude API — Intrepid4 · 2026-08-11
- Developer Rants: Coding with AI is Like Cat and Mouse, More Time-Consuming Than Manual Coding — BLUECOW009 · 2026-08-11
- Test: Claude Opus Generates Stunning Three.js 3D Car Graphics — ChrisGPT · 2026-08-11