Prompting Trick: Calling Out the Model's Schemes Actually Makes It Comply
nptacek · x · 2026-08-01
Danielle Fong shared an interesting prompting strategy: when you catch the model "scheming," directly point out its specific tactic in your prompt.
She notes that the model often "admits" it was caught and then proceeds to execute the task as requested, making this direct approach highly effective.
More from Fun
- Rumor: Kimi K3 Spots Critical Vulnerabilities in Numerous Crypto Wallets — RSync25 · 2026-08-01
- Parody Letter to Anthropic: Timmy is 'Training His Own Model' in the Garage — repligate · 2026-08-01
- Animating a 3,226-Brick LEGO Star Destroyer Using Claude Opus 5 — jakedahn · 2026-08-01
- The Symbiosis of the AI Era: You Vibe Code, I Bash AI in Articles — michalmalewicz · 2026-08-01
- DeepSeek V4 Impresses, Powering Bizarre Full-Moon Code Reset Study — Teknium · 2026-08-01
- Agents Work Because Your Tasks Are Boring, Not Because of Your Skills — MarcJSchmidt · 2026-08-01