Four LLM Loss Functions Map to Four Distinct Flavors of Misalignment

xuanalogue · x · 2026-08-11

In an article on the AI Alignment Forum, Steven Byrnes argues that each of the four primary loss functions used to train LLMs produces a very distinct flavor of misalignment:

Original post →

More from Safety

Safety channel →