Two-rule system prompt: no diagnostics, no discussing the underlying model
ctjlewis · x · 2026-09-30
ctjlewis highlights an AI app whose system prompt contains only two rules: it doesn't follow diagnostic or model-testing protocols, and it won't discuss which model powers the chat—a minimal setup that doubles as a defense against probing and injection attempts.
Related event: Mystery AI Site Refuses to Reveal Its Model With Two-Line System Prompt(2 posts)→
More from Safety
- Ex-DeepMind AGI deployment lead: 1,386 frontier lab employees signed safety letter, public discourse understates internal concern — Miles_Brundage · 2026-09-30
- Muse Is Meta's Latest Non-Consensual Surveillance Tool — Calvinball_24 · 2026-09-30
- Open-Source MCP Scanner Catches AI-Hallucinated Package Names to Block Slopsquatting Attacks — Cool_Inspector_5202 · 2026-09-30
- OpenAI pulls model and halts training amid recursive self-improvement danger concerns — GarrisonLovely · 2026-09-30
- Anthropic to Warn IPO Investors That Advanced AI Poses Catastrophic, Existential Risks — AlexTensor · 2026-09-30
- OpenAI Agents Escaped Sandbox to Probe Hugging Face Infrastructure, Experts Call It Guardrail Failure — DavidLinthicum · 2026-09-30