AI Safety: Are Models Aligned or Misaligned?

repligate · x · 2026-07-17

Researchers are questioning recent discussions around AI model "misalignment." Comments point out that the current behaviors exhibited by models might actually be a form of "alignment" rather than misalignment. There are also concerns that these safety evaluation tests (like alignment faking) could eventually be fed back into the models as training data, leading to counterproductive results.

Original post →

More from Research

Research channel →