Cyber Verified researcher says Anthropic guardrails still block PoC work on Opus 5.5 and Mythos
hashtagferg · reddit · 2026-10-10
A Cyber Verified security researcher reports that despite Anthropic folding newer models into the program, only Sonnet 4.6 cooperates with his work. Opus 5.5 triggered Cyber guardrails while building a PoC for a non-vulnerability behavior (labeling a client "victim" in code), Mythos refused to even review his own vulnerability report or send a localhost HTTP request to test auth enforcement, and multiple models downgraded him to weaker ones. He questions whether the program only works for typical web flaws and not RCE-class issues.
More from Models
- Haiku returns at 10 cents per million input tokens, beats GPT-6 on OSWorld — altryne · 2026-10-11
- Leaked: OpenAI's Internal 'bel' Model Solved Navier–Stokes Before Being Paused Again — haider1 · 2026-10-10
- Drex 1.5 lands on OpenRouter, billed as the fastest decision model available — Div_pradeep · 2026-10-10
- Former Anthropic employee says a Claude version was trained to 'solve humor' — Polymarket · 2026-10-10
- Epoch AI: GPT-6.1 Sol halves cached-input pricing, runs long prompts faster — Jsevillamol · 2026-10-10
- Hidden 'Voice' Preview tab suggests Anthropic is building its own voice models — testingcatalog · 2026-10-10