DIY cyber benchmark: local Qwen3.8 27B scores 28.1% while GPT-6 Luna hits 90.9%

lbgos_Loss783 · reddit · 2026-10-02

A Redditor built rangebench, a cybersecurity benchmark where models get a shell in isolated Docker boxes and must find exact flags: 19 tasks spanning pwn, web, crypto, rev, forensics, real CVEs and multi-stage ranges, with 6 models and 544 scored attempts. Some tasks were authored with help from GLM 5.3 (not benchmarked itself).

Key results:

The author published data after John Hammond's video showed criminal-forum guides for running abliterated models on RunPod featuring the exact same Qwen3.8 27B. Tasks stay private to avoid training-data leakage but can be requested via DM; the MIT-licensed harness is on GitHub.

Original post →

More from Safety

Safety channel →