Can LLMs Compromise LLMs? Model-to-Model Exploitation May Be Underestimated

realsohamparekh · x · 2026-08-28

One underrepresented area in frontier cyber capability evaluations is whether an LLM can compromise another LLM.

While sophisticated benchmarks exist for models exploiting software, and there is separate literature on automated jailbreaking and agent hijacking, these fields should converge. Internal tests show that a correctly prompted smaller model can compromise a more capable agent with better tools or permissions, implying its effective cyber capability may exceed its direct capabilities.

Furthermore, model-to-model exploitation may be trainable and compounding. Fine-tuning relatively small attacker models on successful jailbreak traces could allow them to learn reusable attack patterns and search the space more efficiently than their base capabilities would suggest. This mechanism could also be useful for defense.

Original post →

More from Safety

Safety channel →