Tenants jailbreak shared system prompts — Reddit debates which prompt layout actually stops injection
werunm · reddit · 2026-09-27
A developer on Reddit asked how to stop tenants from writing custom instructions into an assistant's system prompt and switching off scope limits.
Key points from the thread:
- The poster chose moving tenant text down to the user turn as data, reasoning that user turns are data and system turns are instructions — but admits injection regularly lands through user turns too.
- A marked-block layout was seen as the more mature option, though commenters doubted models are obliged to respect such markers.
- No consensus emerged, underscoring that multi-tenant AI products still lack a prompt layout the model will reliably enforce; defense has to come from system-side controls.
More from Safety
- Chatbot picks a user out of a 50-person photo from writing style alone, no face needed — mixy23 · 2026-09-27
- Grok accused of uploading user chat images to the web as Musk says 'this keeps getting worse' — EthanJPerez · 2026-09-27
- $100M industry group accused of funding undisclosed anti-EA attack ads — AaronBergman18 · 2026-09-27
- AI Researcher Dietterich Questions Robotaxi Safety Culture: Too Slow to Fix Problem Behaviors — tdietterich · 2026-09-27
- MikroTik MikroTrick SSH exploit chain PoC goes public, now in CISA KEV — evilsocket · 2026-09-27
- North Carolina Detective Fired for Allegedly Using Flock Cameras to Track a Private Citizen — Polymarket · 2026-09-27