Researchers debate whether training objective matters more than guardrails

sebkrier · x · 2026-07-22

A short exchange about whether a model’s training objective matters more than prompt-level guardrails.

The point being debated is:

It is a compact but substantive discussion about how much alignment should be attributed to training versus runtime prompting.

Related event: Experts Debate Whether Model Failure Constitutes Alignment Issue(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →