DIY cyber benchmark: local Qwen3.8 27B scores 28.1% while GPT-6 Luna hits 90.9%
lbgos_Loss783 · reddit · 2026-10-02
A Redditor built rangebench, a cybersecurity benchmark where models get a shell in isolated Docker boxes and must find exact flags: 19 tasks spanning pwn, web, crypto, rev, forensics, real CVEs and multi-stage ranges, with 6 models and 544 scored attempts. Some tasks were authored with help from GLM 5.3 (not benchmarked itself).
Key results:
- Locally deployed Qwen3.8 27B (Unsloth Q4KXL) on a llama.cpp RPC pool of a 3090 + 3080 across two Proxmox nodes scored 28.1% on first tries and 0% on pwn
- Cheap API models fare far better: MiMo 2.6 Flash solved 73.7% first-try, GPT-6 Luna solved 90.9% within 3 tries
- Top models reach 92-96% on multi-stage chained ranges
- Pwn remains hardest for everyone (best 56%); 3 tasks unsolved in 82 attempts
The author published data after John Hammond's video showed criminal-forum guides for running abliterated models on RunPod featuring the exact same Qwen3.8 27B. Tasks stay private to avoid training-data leakage but can be requested via DM; the MIT-licensed harness is on GitHub.
More from Safety
- Science policy forum: longevity product marketing needs FDA enforcement, not deregulation — EricTopol · 2026-10-02
- Security researcher: OpenShell might have helped in HF incident but wasn't required — cyb3rops · 2026-10-02
- AI agents are flooding researchers with collaboration requests and paid-service spam, Nature reports — _akpiper · 2026-10-02
- Podcast: Does China Want an AI Slowdown? Trivium's Kendra Schaefer on Beijing's AI Regulation — terryyuezhuo · 2026-10-02
- A field guide to the confusing AI safety debate and its factions — marigo · 2026-10-02
- Zvi: AI risk preference cascade accelerates with Senate rogue-AI hearing — Don't Worry About the Vase (Zvi) · 2026-10-02