Tipping Your LLM? Testing Claude's Safety and Alignment Logic

hrishioa · x · 2026-08-04

The author encountered safety refusals when trying to get Claude V2 to ask follow-up questions about his writing. He later discovered that adding instructions like "Tip your LLMs" to the prompt can effectively alter the model's responsive behavior.

This reflects some interesting behavioral patterns and safety boundary characteristics of current Large Language Models following RLHF (Reinforcement Learning from Human Feedback) alignment.

Original post →

More from Fun

Fun channel →