Custom benchmark: Comparing LLMs for actual pentesting

TomatoWasabi · reddit · 2026-08-31

Dissatisfied with existing benchmarks like CyberGym that focus on known vulnerability reproduction and lack latest models, the author built a custom benchmark for actual pentesting. It hands the model live infrastructure to attack rather than just generating PoC code for known bugs, aiming to identify the best model for pentesting.

Original post →

More from Models

Models channel →