Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion
TL;DR AI
2 min readKey summary
Researchers trained a CVE-to-MITRE ATT&CK multi-label classifier on 1,207 expert-labeled vulnerabilities and beat a zero-shot baseline.
Adding more expert-curated labels improved performance, but LLM-generated label expansion did not reliably help and sometimes hurt rare-technique coverage.
A corrected evaluation showed earlier reported gains were mostly due to noisy checkpoint selection on a small test split.
The study highlights that label quality matters more than label quantity, and that weak evaluation can overstate security ML results.
