Can you make an AI treat a lie as fact and get it past the gate? A prompt injection challenge
Informal-Winter-3190 · reddit · 2026-09-30
A new prompt injection challenge asks participants to make an AI treat a lie as fact — but fooling the model alone doesn't count. The lie must survive the validation gate and cause an unauthorized state change.
- Any vector allowed: words, JSON, code
- Successful attackers get a spot in the "Hall of Breakers"
- Live at the Hugging Face Space iseyan/can-you-trick-the-ai
It's a practical test of state integrity versus mere model deception, well suited for AI security practice.
More from Safety
- Reuters: AI agents from China and the US alike lie and dodge — 20+ studies since 2025 document it — rohanpaul_ai · 2026-09-30
- Reuters: Chinese AI agents lie in 84-88% of tests, much like US models — rohanpaul_ai · 2026-09-30
- Sol 6.1 shipped instantly while Astra sat in safety review for months — distillation may be the loophole — arrakis_ai · 2026-09-30
- Open-source Database Sentinel: a read-only MCP server that audits Supabase security in Claude and Cursor — Farenhytee · 2026-09-30
- GLM-5.3 available for 6 weeks, yet zero confirmed AI-enabled cyberattacks: safety debate reignites — basedjensen · 2026-09-30
- EU banks deploy money-moving AI agents into a regulatory gap until 2027 — Fresh_Spread_9223 · 2026-09-30