Courts Sanction Hidden Prompts, but Research Shows Invisible Preference Attacks Work

Robert-Nogacki · reddit · 2026-09-28

A practicing lawyer surveys rulings and research on LLM document manipulation:

The threat ladder: hidden commands (mostly dead, sanctionable) → hidden preferences (effective, detectable) → visible preferences (untested in contracts) → visible signals with no instruction at all, which no filter can address without banning the language of contracts.

Original post →

More from Safety

Safety channel →