Ideal AI safety: Capable but refuses harm

va_joe · x · 2026-08-30

Responds to the query about testing undeployed models on ExploitGym.

Key Point:

Related event: Frontier Models Excel at Attacks but Fail at Defense, Sparking Safety Debate(3 posts)→

Original post →

More from Safety

Safety channel →