CodeBERT — Solidity Vulnerability Classifier (4-class)

Fine-tuned microsoft/codebert-base that classifies a Solidity snippet into one of four vulnerability types: Reentrancy, Integer Overflow, Timestamp Dependency, or Dangerous Delegatecall.

Reproduces the approach of Hossain, Altarawneh & Roberts, "Leveraging LLMs and ML for Smart Contract Vulnerability Detection", IEEE CCWC 2025, on a smaller public dataset.

Code, training pipeline and full evaluation: github.com/Riicko-19/smart-contract-vuln-detector — including a 5-seed variance study, a label-contamination analysis, and an adversarial out-of-distribution probe.

Usage

from transformers import pipeline

clf = pipeline("text-classification", model="AbijithwearsHUGGIES/codebert-smart-contract-vuln")
clf("function withdraw(uint a) public { msg.sender.call.value(a)(\"\"); balances[msg.sender] -= a; }")
# [{'label': 'Reentrancy', 'score': ...}]

Results

This checkpoint, on the held-out 59-contract test split:

Metric Value
Accuracy 0.881
Macro F1 0.871
Reentrancy F1 0.957
Dangerous Delegatecall F1 1.000
Timestamp Dependency F1 0.897
Integer Overflow F1 0.632

Read those with care. The test split has only 59 contracts, so one flipped prediction moves a per-class F1 by roughly 0.09. Measured across 5 seeds, this architecture averages:

Model Accuracy Macro F1 Reentrancy Overflow
CodeBERT 0.868 ±0.037 0.858 ±0.054 0.926 ±0.048 0.624 ±0.170
DistilBERT 0.834 ±0.071 0.826 ±0.061 0.949 ±0.018 0.508 ±0.168

A paired per-seed comparison found no statistically significant difference between CodeBERT and DistilBERT (accuracy p=0.41, macro-F1 p=0.47, Integer Overflow p=0.38). This checkpoint is simply the best single artifact from that sweep, not evidence that CodeBERT is the better architecture here.

Training

  • 12 epochs, class-weighted cross-entropy (the dataset is imbalanced)
  • Best checkpoint selected on validation macro-F1, not eval_loss — selecting on loss lets the majority class dominate and yields 0.00 F1 on Integer Overflow, because the "best" checkpoint abandons the minority class
  • Max sequence length 512 (no truncation occurs; longest contract is 248 tokens)

Data

387 labelled contracts from the Messi-Q/Smart-Contract-Dataset "Resource 2" release — Timestamp Dependency 174, Integer Overflow 80, Reentrancy 71, Dangerous Delegatecall 62. Split 270 train / 58 val / 59 test.

This is roughly a sixth the size of the paper's 2,217-contract IR-Fuzz split (which is not publicly redistributable), so treat these numbers as directional rather than a reproduction.

Limitations

  • Integer Overflow is weak (F1 ~0.62, ±0.17 across seeds), and the cause is the label, not the model. Across all 387 contracts, the other three classes have a near-perfect syntactic signature (Reentrancy 100% contain .call.value(, Timestamp 100% contain block.timestamp/now, Delegatecall 89% contain .delegatecall(). Integer Overflow has none of its own: 39% of its contracts contain the reentrancy pattern and 55% contain timestamp calls. The source dataset ships four independent binary labelled sets, and forcing a single label onto contracts that exhibit several vulnerabilities pushes the ambiguity into this class. Treating the task as multi-label would be the principled fix.
  • Forced-choice, not detection. The model always returns one of four classes. It cannot say "no vulnerability", and it was trained only on vulnerable contracts. An out-of-distribution probe makes the cost concrete: a pure math library with no state, no external calls and no timestamps is classified Timestamp Dependency at 99.1% confidence, and Python source code scores Integer Overflow at 93.7%. Confidence does not help — on benign inputs it reaches 99.1%, while genuine vulnerabilities go as low as 97.9%, so the ranges overlap and no threshold separates them. Never use this to decide whether a contract is safe.
  • It does generalise on genuinely vulnerable code, which is the flip side: it correctly flags reentrancy written with modern .call{value:} syntax, and even reentrancy expressed through a callback with no low-level call at all — so it is not merely keyword matching.
  • Not a substitute for an audit. This is a research/portfolio artifact trained on 387 contracts, not a security tool.

License

MIT

Downloads last month
64
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AbijithwearsHUGGIES/codebert-smart-contract-vuln

Quantized
(6)
this model

Space using AbijithwearsHUGGIES/codebert-smart-contract-vuln 1

Paper for AbijithwearsHUGGIES/codebert-smart-contract-vuln