mitsuhiko on LLM Prefilling: Injecting Fake Assistant Messages Risks Safety

mitsuhiko · x · 2026-08-12

Discussing whether closed-source SOTA model labs will ban assistant prefilling in newer models, prominent developer Armin Ronacher (@mitsuhiko) pointed out a core issue: developers can inject fabricated content into the transcript as 'assistant messages' that were never actually generated by the LLM.

While this mechanism provides flexibility for application development, it can easily be exploited to bypass model safety guardrails or manipulate context, suggesting it might face stricter platform restrictions in the future.

Related event: Developer Mitsuhiko Warns of Security Risks in LLM Assistant Prefilling(2 posts)→

Original post →

More from Models

Models channel →