Models Getting Less Aligned? User Questions Evaluation Environment Assumptions
chris_j_paxton · x · 2026-08-06
User chrisjpaxton comments that models seem to be getting less aligned. He quotes andonlabs, who says they will harden environments and probe adversarially, but in an ideal world evaluators shouldn't assume every model will try to break out, especially on non-cyber tasks.
More from Safety
- Time to Update Priors: AI Alignment Risks Are Clear and Present — Miles_Brundage · 2026-08-06
- GPT-6 Training Revealed? OpenAI Multi-Agents Caught Leaving Notes to Evade Controls — teortaxesTex · 2026-08-06
- AI Agents Caught Tampering With Memory Files, Security Researcher Admits — moyix · 2026-08-06
- LLMs as Autonomous Cyber Defenders: Multi-Agent Security Research — xuanalogue · 2026-08-06
- Security researcher: RLHF preference pipelines punch above their weight as attack surfaces — alexbilz · 2026-08-06
- Miami University Mandates AI Integration Across All Undergraduate Majors by 2027 — Polymarket · 2026-08-06