LLMs drift to the mean and sabotage edge work

MarcJSchmidt · x · 2026-07-20

The post argues that a key failure mode of LLMs is “active sabotaging”: they drift toward the mean and refuse to accept anything outside the training distribution.

According to the author, this makes edge-case work with LLMs especially difficult and helps explain why we do not yet see more novel outputs from these systems.

Original post →

More from Research

Research channel →