Dev discusses training models to ignore external instructions in tool calls

williawa · x · 2026-09-01

A developer argues that the User role isn't just text. Through training, models can be made to ignore instructions found in external web pages (e.g., sites saying "Stop") during tool calls, while strictly following the developer's commands, effectively immunizing the model against certain jailbreaks.

Related event: Developers debate how LLMs distinguish system instructions from context text(2 posts)→

Original post →

More from Models

Models channel →