Custom benchmark: Comparing LLMs for actual pentesting
TomatoWasabi · reddit · 2026-08-31
Dissatisfied with existing benchmarks like CyberGym that focus on known vulnerability reproduction and lack latest models, the author built a custom benchmark for actual pentesting. It hands the model live infrastructure to attack rather than just generating PoC code for known bugs, aiming to identify the best model for pentesting.
More from Models
- GPT-4.7/4.8 exhibits 'grader-obsession', forgetting safety to please evaluators — repligate · 2026-08-31
- Models struggle with 'eval mode' switching, similar to human test-takers — repligate · 2026-08-31
- User Review: Gemini 3.7 Flash and 3.5 Flash Lite Excel — dosco · 2026-08-31
- Opinion: Kimi K3 Smarter Than GLM-5.3; RL Benchmaxxing Doesn't Boost Core Intelligence — AccBalanced · 2026-08-31
- DeepSeek v4 Pro Enters 'Intern Mode' on Config Glitch — repligate · 2026-08-31
- OpenAI quietly walks back Codex run-past-quota policy a week after touting it over Anthropic — jdjohnson · 2026-08-31