KaliBench: 8,504 pairs benchmark shows no open-weight LLM exceeds 42% on Kali Linux CLI tasks

RISys-Lab · hf · 2026-10-02

KaliBench is a fine-grained benchmark for translating natural language into executable Kali Linux cybersecurity CLI commands: 8,504 query-command pairs across 1,642 tools, 23 capability dimensions and 5 security phases, with deterministic canonicalization and alias-aware evaluation. Across 24 configurations of open-weight models, none exceeds 42% exact-command accuracy without tool hints. The benchmark also provides runtime-free verifiable rewards: SFT + RLVR lifts an 8B model to parity with a 685B MoE.

Original post →

More from Safety

Safety channel →