LLM-generated CVE labels fail to improve ATT&CK mapping, study finds
CIRCL · hf · 2026-07-29
Key points
CIRCL presents a reproducible pipeline for mapping free-text CVE descriptions to MITRE ATT&CK Enterprise techniques.
- The model is trained on a curated gold set of 1,207 CVEs from expert MITRE CTID mappings.
- It roughly doubles recall@5 versus a zero-shot embedding-similarity baseline and improves every ranking metric.
- The authors test whether LLM-assisted label expansion can grow the dataset, but find the apparent gains are an evaluation artifact.
- Independent replication and a larger expansion study show LLM-generated labels do not reliably improve performance, with about 0.39 agreement with expert annotations and weaker rare-technique coverage at larger scale.
- The real bottleneck is label quality, not dataset size: adding expert-curated data helps consistently, while LLM-labeled data does not.
- All datasets, models, code, and logs are released publicly.
More from Safety
- Hugging Face says it used an open model to defend against an autonomous agent cyberattack — max_paperclips · 2026-07-29
- Anthropic copyright ruling sparks debate over book destruction and superintelligent lawyers — AndyMasley · 2026-07-29
- EU AI Act rolls out with risk-based rules and bans on clearly harmful practices — emmanuelvivier · 2026-07-29
- Is AI a New Form of IP? Industry Debates Open Weights vs. Ownership — aryaman2020 · 2026-07-29
- John H. Cochrane blasts the rush to regulate AI in a new essay — Dan_Jeffries1 · 2026-07-29
- AI slowing is no longer fringe as labs and doom debates collide — basedjensen · 2026-07-29