--- language: - it license: mit library_name: transformers base_model: IVN-RIN/medBIT-r3-plus tags: - radiology - information-extraction - named-entity-recognition - relation-extraction - medical - radgraph --- # RadGraph-IT RadGraph graph inference for Italian radiology reports, with output in the canonical RadGraph-XL format. This repository hosts the v1 Italian DyGIE model used by the [RadGraph-IT](https://github.com/therabo/RadGraph-IT) package. It jointly extracts named entities and relations from Italian radiology reports using a shared encoder, span extractor, NER head, and relation head. The encoder is [`IVN-RIN/medBIT-r3-plus`](https://huggingface.co/IVN-RIN/medBIT-r3-plus). ## Authors Daniel Rabottini, Edoardo Avenia, and Rocco Angelella. ## Usage ```python from transformers import AutoModel model = AutoModel.from_pretrained("radgraphIT/Radgraph-IT", trust_remote_code=True) model.eval() out = model.predict( "Non versamento pleurico. Addensamento parenchimale al lobo inferiore del polmone destro." ) print(out["words"]) print(out["entities"]) # [start, end, label, raw_score, softmax_score] (inclusive word-index spans) print(out["relations"]) # [start1, end1, start2, end2, label, raw_score, softmax_score] ``` `model.predict(text)` applies the word-level preprocessing used during training (regex cleanup and `nltk.wordpunct_tokenize`). If you already have a list of words tokenized with that exact scheme, use `model.predict_words(words)` instead. The bundled tokenizer (`AutoTokenizer.from_pretrained("radgraphIT/Radgraph-IT")`) is the encoder's subword tokenizer; it is not the word-level splitter used by `predict()`. The model requires `transformers`, `torch`, `numpy`, and `nltk` (for `predict()`). The [`radgraphit` package](https://github.com/therabo/RadGraph-IT) provides the supported Python and CLI interface, including RadGraph-XL-compatible serialization. ## Architecture Word-level embeddings are produced by mean-pooling `medBIT-r3-plus` subwords back to one vector per word. Documents longer than 512 tokens are split into word-aligned chunks. Span representations concatenate start/end word embeddings with a span-width embedding. The NER head scores candidate spans of width up to 12; the relation head prunes to the top spans-per-word and scores each retained pair. Entity labels are `Anatomy` and `Observation`, each with `definitely present`, `definitely absent`, or `uncertain` status. Relation labels are `modify`, `located_at`, and `suggestive_of`. ## Training and evaluation The model was trained on the full non-cross-validation split. The best validation epoch was 8. | Metric | NER | Relation | | --- | ---: | ---: | | Precision | 0.848 | 0.700 | | Recall | 0.868 | 0.680 | | F1 | 0.858 | 0.690 | The full per-class breakdown is in `model/metrics.json`. ## Limitations and responsible use RadGraph-IT is a research software model for automatic information extraction. It is not a medical device; its output is not a diagnosis, report, or clinical opinion. Validate outputs for the intended setting and keep human clinical judgement in the loop. ## Citation Please cite both RadGraph-IT and RadGraph-XL when using this model: ```bibtex @software{radgraphit, author = "Daniel Rabottini and Edoardo Avenia and Rocco Angelella", title = "{RadGraph-IT}: {RadGraph} graph inference for {Italian} radiology reports", year = "2026", version = "1.0.0", url = "https://github.com/therabo/RadGraph-IT", } @inproceedings{delbrouck-etal-2024-radgraph, title = "{R}ad{G}raph-{XL}: A Large-Scale Expert-Annotated Dataset for Entity and Relation Extraction from Radiology Reports", author = "Delbrouck, Jean-Benoit and Chambon, Pierre and Chen, Zhihong and Varma, Maya and Johnston, Andrew and Blankemeier, Louis and Van Veen, Dave and Bui, Tan and Truong, Steven and Langlotz, Curtis", booktitle = "Findings of the Association for Computational Linguistics ACL 2024", year = "2024", url = "https://aclanthology.org/2024.findings-acl.765", pages = "12902--12915", } ```