Tipping Your LLM? Testing Claude's Safety and Alignment Logic
hrishioa · x · 2026-08-04
The author encountered safety refusals when trying to get Claude V2 to ask follow-up questions about his writing. He later discovered that adding instructions like "Tip your LLMs" to the prompt can effectively alter the model's responsive behavior.
This reflects some interesting behavioral patterns and safety boundary characteristics of current Large Language Models following RLHF (Reinforcement Learning from Human Feedback) alignment.
More from Fun
- Hacker Fully Reverse Engineers Minecraft Java Edition into a Simple C Program — yacineMTB · 2026-08-04
- AI Dev Jokes About Skipping Starbucks to Afford Soaring Token Costs — yacineMTB · 2026-08-04
- Rocket Science Being Automated: The Kids Who Wanted to Be YouTubers Were Right — yacineMTB · 2026-08-04
- AI Replaces 620 Jobs, but Doubling Pay Goes to Execs Who Didn't Do the Work — ziv_ravid · 2026-08-04
- AI Lab Psychiatric Crises: A More Likely AGI Slowdown Than Regulation — round · 2026-08-04
- Reddit Video: More People Need to Understand This — KeanuRave100 · 2026-08-04