KaliBench: 8,504 pairs benchmark shows no open-weight LLM exceeds 42% on Kali Linux CLI tasks
RISys-Lab · hf · 2026-10-02
KaliBench is a fine-grained benchmark for translating natural language into executable Kali Linux cybersecurity CLI commands: 8,504 query-command pairs across 1,642 tools, 23 capability dimensions and 5 security phases, with deterministic canonicalization and alias-aware evaluation. Across 24 configurations of open-weight models, none exceeds 42% exact-command accuracy without tool hints. The benchmark also provides runtime-free verifiable rewards: SFT + RLVR lifts an 8B model to parity with a 685B MoE.
More from Safety
- TA419 hackers impersonate ex-White House AI official Lynne Parker to phish AI policy experts — cyb3rops · 2026-10-02
- Reddit asks: when AI denies you service, who do you argue with? — yi111 · 2026-10-02
- Six months after Anthropic's 'too dangerous to release' Mythos, the predicted vulnerability explosion never came — redbaron_4 · 2026-10-02
- Cantina releases apex-flash-1, an open-weights security model post-trained on real paid vulnerabilities — ctjlewis · 2026-10-02
- OpenRouter launches Security Center after finding 1,000+ dormant API keys across 85 employees — AccBalanced · 2026-10-02
- Open weights called "dangerous" — echoing every tech that shifted information power — 0xAllen_ · 2026-10-02