Cyber Verified researcher says Anthropic guardrails still block PoC work on Opus 5.5 and Mythos

hashtagferg · reddit · 2026-10-10

A Cyber Verified security researcher reports that despite Anthropic folding newer models into the program, only Sonnet 4.6 cooperates with his work. Opus 5.5 triggered Cyber guardrails while building a PoC for a non-vulnerability behavior (labeling a client "victim" in code), Mythos refused to even review his own vulnerability report or send a localhost HTTP request to test auth enforcement, and multiple models downgraded him to weaker ones. He questions whether the program only works for typical web flaws and not RCE-class issues.

Original post →

More from Models

Models channel →