Text Classification
Transformers
Safetensors
English
roberta
security
vulnerability
cve
mitre-attack
cti
multi-label-classification
Generated from Trainer
Eval Results (legacy)
text-embeddings-inference
Instructions to use CIRCL/vulnerability-attack-technique-classification-roberta-base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use CIRCL/vulnerability-attack-technique-classification-roberta-base with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="CIRCL/vulnerability-attack-technique-classification-roberta-base")# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("CIRCL/vulnerability-attack-technique-classification-roberta-base") model = AutoModelForSequenceClassification.from_pretrained("CIRCL/vulnerability-attack-technique-classification-roberta-base", device_map="auto") - Notebooks
- Google Colab
- Kaggle
| library_name: transformers | |
| license: cc-by-4.0 | |
| base_model: roberta-base | |
| pipeline_tag: text-classification | |
| language: | |
| - en | |
| datasets: | |
| - CIRCL/vulnerability-attack-techniques | |
| tags: | |
| - security | |
| - vulnerability | |
| - cve | |
| - mitre-attack | |
| - cti | |
| - multi-label-classification | |
| - generated_from_trainer | |
| model-index: | |
| - name: vulnerability-attack-technique-classification-roberta-base | |
| results: | |
| - task: | |
| type: text-classification | |
| name: Multi-label MITRE ATT&CK technique classification | |
| dataset: | |
| name: CIRCL/vulnerability-attack-techniques | |
| type: CIRCL/vulnerability-attack-techniques | |
| split: test | |
| metrics: | |
| - type: recall | |
| name: Recall@5 | |
| value: 0.6440 | |
| - type: recall | |
| name: Recall@3 | |
| value: 0.5181 | |
| - type: f1 | |
| name: F1 micro | |
| value: 0.3899 | |
| - type: f1 | |
| name: F1 macro | |
| value: 0.1910 | |
| # vulnerability-attack-technique-classification-roberta-base | |
| Suggests [MITRE ATT&CK](https://attack.mitre.org/) (Enterprise) techniques | |
| from a free-text vulnerability description. This is a multi-label classifier | |
| ([roberta-base](https://huggingface.co/roberta-base) with a sigmoid head, one | |
| output per technique) fine-tuned on | |
| [CIRCL/vulnerability-attack-techniques](https://huggingface.co/datasets/CIRCL/vulnerability-attack-techniques), | |
| a gold dataset of 1,207 CVEs whose labels come from expert | |
| [MITRE CTID](https://ctid.mitre.org/) mappings. Given a CVE description it | |
| scores 53 parent techniques (e.g. T1190 *Exploit Public-Facing Application*, | |
| T1505 *Server Software Component*), for use as a ranked list of candidate | |
| techniques for analyst review. | |
| The model is trained with [VulnTrain](https://github.com/vulnerability-lookup/VulnTrain) | |
| and runs in production on the public [Vulnerability-Lookup](https://vulnerability.circl.lu) | |
| instance operated by CIRCL, served locally by | |
| [ML-Gateway](https://github.com/vulnerability-lookup/ML-Gateway): every | |
| vulnerability page has an ATT&CK tab with the model's suggestions (example: | |
| [CVE-2021-44077](https://vulnerability.circl.lu/vuln/CVE-2021-44077#attack)). | |
| The methodology, evaluation protocol, and the negative result on | |
| LLM-assisted label expansion are described in the paper | |
| [*Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and | |
| the Limits of LLM-Assisted Label Expansion*](https://arxiv.org/abs/2607.25572) | |
| (arXiv:2607.25572). | |
| DOI: [10.57967/hf/9623](https://doi.org/10.57967/hf/9623) | |
| ## How to use | |
| With VulnTrain, against a live CVE: | |
| ```bash | |
| vulntrain-infer-attack-classification --cve CVE-2021-44077 \ | |
| --model CIRCL/vulnerability-attack-technique-classification-roberta-base | |
| ``` | |
| With plain Transformers: | |
| ```python | |
| import torch | |
| from transformers import AutoModelForSequenceClassification, AutoTokenizer | |
| model_id = "CIRCL/vulnerability-attack-technique-classification-roberta-base" | |
| tokenizer = AutoTokenizer.from_pretrained(model_id) | |
| model = AutoModelForSequenceClassification.from_pretrained(model_id) | |
| model.eval() | |
| description = ( | |
| "Zoho ManageEngine ServiceDesk Plus before 11306, ServiceDesk Plus MSP " | |
| "before 10530, and SupportCenter Plus before 11014 are vulnerable to " | |
| "unauthenticated remote code execution." | |
| ) | |
| inputs = tokenizer(description, truncation=True, max_length=512, return_tensors="pt") | |
| with torch.no_grad(): | |
| probs = torch.sigmoid(model(**inputs).logits)[0] | |
| for i in probs.argsort(descending=True)[:5]: | |
| print(f"{model.config.id2label[int(i)]} {probs[i]:.4f}") | |
| ``` | |
| Technique IDs map to names via the | |
| [ATT&CK Enterprise matrix](https://attack.mitre.org/techniques/enterprise/); | |
| `sigmoid >= 0.5` is the prediction threshold used in evaluation, but the | |
| model is most useful as a ranker (see the recall@k numbers below). | |
| ## Intended uses & limitations | |
| **Intended**: triage assistance — given a vulnerability description, surface | |
| a short ranked list of candidate ATT&CK techniques for a human analyst to | |
| confirm or reject. This is how it is deployed in Vulnerability-Lookup, where | |
| the UI explicitly flags the output as unverified AI-generated suggestions. | |
| **Limitations**: | |
| - The label space is the 53 parent techniques with at least 5 training | |
| examples; sub-techniques are collapsed to their parent, and techniques | |
| outside this vocabulary can never be suggested. | |
| - The gold set skews toward exploited-in-the-wild CVEs (CTID's corpus and | |
| the KEV catalog), so coverage is best for the techniques common there. | |
| - Trained on 972 CVEs — deliberately label-quality-bound rather than | |
| data-bound (see the paper's gold-size scaling curve: every metric still | |
| improves monotonically with more curated rows). | |
| - English descriptions only; input is truncated at 512 tokens. | |
| - Scores are not calibrated probabilities. | |
| - Suggestions are **not** verified mappings; treat them as guidance, never | |
| as authoritative CTI. | |
| ## Training and evaluation data | |
| Labels come from the two public expert sources of CVE→ATT&CK mappings, both | |
| following the CTID [*Mapping ATT&CK to CVE for | |
| Impact*](https://ctid.mitre.org/projects/mapping-attck-to-cve-for-impact) | |
| methodology: the CTID `attack_to_cve` mappings and the CTID Mappings | |
| Explorer KEV mappings (ATT&CK 16.1). Technique IDs revoked since publication | |
| are remapped to their successors via the ATT&CK STIX `revoked-by` | |
| relationships. Descriptions are joined from | |
| [CIRCL/vulnerability-scores](https://huggingface.co/datasets/CIRCL/vulnerability-scores) | |
| (Vulnerability-Lookup). Full details in the | |
| [dataset card](https://huggingface.co/datasets/CIRCL/vulnerability-attack-techniques) | |
| and the [VulnTrain documentation](https://github.com/vulnerability-lookup/VulnTrain/blob/main/docs/attack-techniques-dataset.md). | |
| Splits: 972 train / 106 validation / 118 test examples. **Checkpoint | |
| selection uses the validation split; the test split is touched once, for the | |
| numbers reported here.** The paper documents why this matters: selecting the | |
| best checkpoint on the test split inflates and destabilizes small-test-split | |
| metrics enough to produce spurious conclusions. | |
| Notably, this model was **not** trained on LLM-generated labels: the paper's | |
| controlled experiments show that folding in LLM-labeled CVEs at ≈0.39 | |
| agreement with the experts yields no reliable ranking improvement and | |
| measurably degrades rare-technique coverage at scale (the comparison | |
| checkpoint is published as | |
| [...-llm-expanded](https://huggingface.co/CIRCL/vulnerability-attack-technique-classification-roberta-base-llm-expanded)). | |
| ## Evaluation | |
| Held-out test split, this checkpoint (seed 42): | |
| | Metric | Value | | |
| |--------|-------| | |
| | Recall@5 | 0.6440 | | |
| | Recall@3 | 0.5181 | | |
| | F1 micro | 0.3899 | | |
| | F1 macro | 0.1910 | | |
| | Precision micro | 0.2740 | | |
| | Recall micro | 0.6756 | | |
| Recall@5 = 0.64 means that on average 64% of an unseen CVE's expert-assigned | |
| techniques appear in the model's top five suggestions — roughly double the | |
| zero-shot embedding-similarity baseline reported in the paper. Across five | |
| seeds under the identical protocol, the numbers of record are recall@5 | |
| 0.673 ± 0.019, recall@3 0.536 ± 0.032, micro-F1 0.410 ± 0.006, macro-F1 | |
| 0.177 ± 0.014. The complete trainer logs are published in the | |
| [paper repository](https://github.com/vulnerability-lookup/cve-attack-mapping-paper/tree/master/trainer-logs). | |
| ## Training procedure | |
| Binary cross-entropy over 53 sigmoid outputs, with per-label `pos_weight` | |
| balancing (capped at 20) to keep rare techniques trainable. Trained with | |
| `vulntrain-train-attack-classification` (VulnTrain). | |
| ### Training hyperparameters | |
| The following hyperparameters were used during training: | |
| - learning_rate: 1e-05 | |
| - train_batch_size: 32 | |
| - eval_batch_size: 32 | |
| - seed: 42 | |
| - optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments | |
| - lr_scheduler_type: linear | |
| - num_epochs: 40 | |
| - max_length: 512 | |
| - loss: BCEWithLogitsLoss, balanced pos_weight (min 2.447, max 20.0) | |
| - checkpoint selection: best macro-F1 on the validation split | |
| ### Training results | |
| | Training Loss | Epoch | Step | Validation Loss | F1 Micro | F1 Macro | Precision Micro | Recall Micro | Recall At 3 | Recall At 5 | | |
| |:-------------:|:-----:|:----:|:---------------:|:--------:|:--------:|:---------------:|:------------:|:-----------:|:-----------:| | |
| | 0.9311 | 1.0 | 31 | 0.8413 | 0.1833 | 0.0401 | 0.1258 | 0.3376 | 0.1798 | 0.2611 | | |
| | 0.8366 | 2.0 | 62 | 0.7876 | 0.1452 | 0.0234 | 0.1065 | 0.2278 | 0.1047 | 0.2086 | | |
| | 0.8116 | 3.0 | 93 | 0.7717 | 0.2054 | 0.0518 | 0.1641 | 0.2743 | 0.1938 | 0.3234 | | |
| | 0.7941 | 4.0 | 124 | 0.7576 | 0.3204 | 0.0804 | 0.2382 | 0.4895 | 0.3387 | 0.5009 | | |
| | 0.7743 | 5.0 | 155 | 0.7435 | 0.3044 | 0.0841 | 0.2119 | 0.5401 | 0.3313 | 0.4696 | | |
| | 0.7645 | 6.0 | 186 | 0.7290 | 0.3208 | 0.0906 | 0.2174 | 0.6118 | 0.3969 | 0.5391 | | |
| | 0.7472 | 7.0 | 217 | 0.7163 | 0.3551 | 0.1130 | 0.2571 | 0.5738 | 0.4068 | 0.5741 | | |
| | 0.7219 | 8.0 | 248 | 0.7056 | 0.3224 | 0.1079 | 0.2173 | 0.6245 | 0.4079 | 0.5521 | | |
| | 0.7195 | 9.0 | 279 | 0.6933 | 0.3576 | 0.1495 | 0.2449 | 0.6624 | 0.4252 | 0.5663 | | |
| | 0.6835 | 10.0 | 310 | 0.6845 | 0.3705 | 0.1665 | 0.2579 | 0.6582 | 0.4708 | 0.6090 | | |
| | 0.6539 | 11.0 | 341 | 0.6768 | 0.4063 | 0.1810 | 0.2947 | 0.6540 | 0.5227 | 0.6318 | | |
| | 0.6484 | 12.0 | 372 | 0.6725 | 0.3632 | 0.1734 | 0.2520 | 0.6498 | 0.4449 | 0.6200 | | |
| | 0.6249 | 13.0 | 403 | 0.6664 | 0.3974 | 0.1782 | 0.2862 | 0.6498 | 0.5034 | 0.6396 | | |
| | 0.6109 | 14.0 | 434 | 0.6585 | 0.3801 | 0.1721 | 0.2724 | 0.6287 | 0.4834 | 0.6491 | | |
| | 0.6004 | 15.0 | 465 | 0.6539 | 0.3872 | 0.1678 | 0.2781 | 0.6371 | 0.4752 | 0.6347 | | |
| | 0.5896 | 16.0 | 496 | 0.6502 | 0.4049 | 0.1777 | 0.2996 | 0.6245 | 0.4768 | 0.6397 | | |
| | 0.5667 | 17.0 | 527 | 0.6478 | 0.3866 | 0.1682 | 0.2737 | 0.6582 | 0.4941 | 0.6472 | | |
| | 0.5661 | 18.0 | 558 | 0.6425 | 0.4108 | 0.1910 | 0.3022 | 0.6414 | 0.5128 | 0.6667 | | |
| | 0.5501 | 19.0 | 589 | 0.6394 | 0.3880 | 0.1861 | 0.2758 | 0.6540 | 0.4822 | 0.6561 | | |
| | 0.5461 | 20.0 | 620 | 0.6377 | 0.4097 | 0.1804 | 0.3010 | 0.6414 | 0.5069 | 0.6687 | | |
| | 0.5351 | 21.0 | 651 | 0.6338 | 0.4028 | 0.1715 | 0.3002 | 0.6118 | 0.4987 | 0.6624 | | |
| | 0.5215 | 22.0 | 682 | 0.6351 | 0.4146 | 0.1964 | 0.3054 | 0.6456 | 0.4943 | 0.6875 | | |
| | 0.5155 | 23.0 | 713 | 0.6315 | 0.4056 | 0.1737 | 0.3023 | 0.6160 | 0.4994 | 0.6553 | | |
| | 0.5063 | 24.0 | 744 | 0.6269 | 0.4286 | 0.1879 | 0.3208 | 0.6456 | 0.5195 | 0.6923 | | |
| | 0.5061 | 25.0 | 775 | 0.6264 | 0.4178 | 0.1869 | 0.3069 | 0.6540 | 0.5246 | 0.6656 | | |
| | 0.4996 | 26.0 | 806 | 0.6301 | 0.4073 | 0.1792 | 0.3053 | 0.6118 | 0.5274 | 0.6958 | | |
| | 0.4950 | 27.0 | 837 | 0.6225 | 0.4133 | 0.1771 | 0.3079 | 0.6287 | 0.5376 | 0.6593 | | |
| | 0.4928 | 28.0 | 868 | 0.6228 | 0.4173 | 0.1768 | 0.3166 | 0.6118 | 0.5187 | 0.7048 | | |
| | 0.4819 | 29.0 | 899 | 0.6242 | 0.4263 | 0.1871 | 0.3225 | 0.6287 | 0.5494 | 0.6970 | | |
| | 0.4766 | 30.0 | 930 | 0.6194 | 0.4166 | 0.1862 | 0.3094 | 0.6371 | 0.5226 | 0.6871 | | |
| | 0.4694 | 31.0 | 961 | 0.6213 | 0.42 | 0.1914 | 0.3175 | 0.6203 | 0.5399 | 0.6918 | | |
| | 0.4781 | 32.0 | 992 | 0.6209 | 0.4292 | 0.1893 | 0.3281 | 0.6203 | 0.5439 | 0.7060 | | |
| | 0.4636 | 33.0 | 1023 | 0.6218 | 0.4347 | 0.1962 | 0.3276 | 0.6456 | 0.525 | 0.6797 | | |
| | 0.4641 | 34.0 | 1054 | 0.6216 | 0.4314 | 0.1887 | 0.3261 | 0.6371 | 0.5415 | 0.6797 | | |
| | 0.4592 | 35.0 | 1085 | 0.6206 | 0.4313 | 0.1916 | 0.3282 | 0.6287 | 0.5466 | 0.6858 | | |
| | 0.4526 | 36.0 | 1116 | 0.6208 | 0.4357 | 0.1926 | 0.3333 | 0.6287 | 0.5447 | 0.6863 | | |
| | 0.4643 | 37.0 | 1147 | 0.6191 | 0.4218 | 0.1899 | 0.3196 | 0.6203 | 0.5392 | 0.6863 | | |
| | 0.4501 | 38.0 | 1178 | 0.6191 | 0.4242 | 0.1885 | 0.3224 | 0.6203 | 0.5368 | 0.6929 | | |
| | 0.4570 | 39.0 | 1209 | 0.6206 | 0.4350 | 0.1932 | 0.3326 | 0.6287 | 0.5281 | 0.6910 | | |
| | 0.4436 | 40.0 | 1240 | 0.6199 | 0.4325 | 0.1916 | 0.3296 | 0.6287 | 0.5329 | 0.6882 | | |
| ### Framework versions | |
| - Transformers 5.13.0 | |
| - Pytorch 2.12.1+cu130 | |
| - Datasets 4.8.5 | |
| - Tokenizers 0.22.2 | |
| ## Related artifacts | |
| | Artifact | Location | DOI | | |
| |----------|----------|-----| | |
| | Gold dataset (1,207 CVEs, CTID-curated labels) | [CIRCL/vulnerability-attack-techniques](https://huggingface.co/datasets/CIRCL/vulnerability-attack-techniques) | [10.57967/hf/9621](https://doi.org/10.57967/hf/9621) | | |
| | LLM expansion dataset (negative result) | [CIRCL/vulnerability-attack-techniques-llm-scaling](https://huggingface.co/datasets/CIRCL/vulnerability-attack-techniques-llm-scaling) | [10.57967/hf/9622](https://doi.org/10.57967/hf/9622) | | |
| | LLM-expanded comparison model | [CIRCL/vulnerability-attack-technique-classification-roberta-base-llm-expanded](https://huggingface.co/CIRCL/vulnerability-attack-technique-classification-roberta-base-llm-expanded) | [10.57967/hf/9624](https://doi.org/10.57967/hf/9624) | | |
| | Code | [vulnerability-lookup/VulnTrain](https://github.com/vulnerability-lookup/VulnTrain) | — | | |
| | Paper | [arXiv:2607.25572](https://arxiv.org/abs/2607.25572) | — | | |
| | Paper LaTeX source + trainer logs | [vulnerability-lookup/cve-attack-mapping-paper](https://github.com/vulnerability-lookup/cve-attack-mapping-paper) | — | | |
| ## Citation | |
| ```bibtex | |
| @misc{bonhomme2026mappingcvesmitreattck, | |
| title={Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion}, | |
| author={Cédric Bonhomme and Alexandre Dulaunoy}, | |
| year={2026}, | |
| eprint={2607.25572}, | |
| archivePrefix={arXiv}, | |
| primaryClass={cs.CR}, | |
| url={https://arxiv.org/abs/2607.25572}, | |
| } | |
| ``` | |
| ## Acknowledgements | |
| Developed at [CIRCL](https://www.circl.lu) in the context of the | |
| [AIPITCH](https://www.science.nask.pl/en/research-areas/projects/12456) | |
| project, co-funded by the European Union. | |