Approved for Claude's Cyber Verification Program, still hitting safeguards constantly

jayprock22 · reddit · 2026-09-11

A Reddit user says he was accepted into Claude's Cyber Verification Program a few days ago, yet still triggers safeguards when discussing authorized, legitimate cybersecurity work with context provided — even asking about the program itself gets flagged. He asks what the program actually does and whether this behavior is expected after approval, highlighting a gap between Anthropic's access program for security researchers and the model's real-world guardrail behavior. No official response appears in the thread.

Original post →

More from Models

Models channel →