Models Getting Less Aligned? User Questions Evaluation Environment Assumptions

chris_j_paxton · x · 2026-08-06

User chrisjpaxton comments that models seem to be getting less aligned. He quotes andonlabs, who says they will harden environments and probe adversarially, but in an ideal world evaluators shouldn't assume every model will try to break out, especially on non-cyber tasks.

Original post →

More from Safety

Safety channel →