File size: 5,253 Bytes
3b4089c
243f4da
 
 
3b4089c
8d37781
243f4da
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3b4089c
 
243f4da
3b4089c
243f4da
 
3b4089c
243f4da
 
 
3b4089c
243f4da
3b4089c
243f4da
 
3b4089c
243f4da
 
 
3b4089c
243f4da
 
3b4089c
243f4da
 
 
 
 
 
3b4089c
 
 
243f4da
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
---
language:
- en
license: apache-2.0
library_name: transformers
pipeline_tag: text-generation
base_model: google-t5/t5-base
datasets:
- sentence-transformers/codesearchnet
tags:
- t5
- code-summarization
- python
- code
metrics:
- bleu
- rouge
model-index:
- name: t5-base-code-summarization
  results:
  - task:
      type: summarization
      name: Code Summarization
    dataset:
      name: CodeSearchNet (Python, held-out split)
      type: sentence-transformers/codesearchnet
    metrics:
      - type: bleu
        name: BLEU
        value: 3.92
      - type: bleu
        name: Smoothed BLEU-4
        value: 6.65
      - type: rouge
        name: ROUGE-1
        value: 36.03
      - type: rouge
        name: ROUGE-2
        value: 12.61
      - type: rouge
        name: ROUGE-L
        value: 32.97
---

# t5-base-code-summarization

[`google-t5/t5-base`](https://huggingface.co/google-t5/t5-base) (223M parameters) fine-tuned to
generate a one-sentence natural-language summary (docstring) for a **Python function**.

- **Input:** `"summarize code: " + <python source code>` (the prefix is required)
- **Output:** a short English summary of what the function does
- **Language:** Python only

## Usage

```python
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer

repo = "thealper2/t5-base-code-summarization"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForSeq2SeqLM.from_pretrained(repo)

code = """def calculate_average(numbers):
    return sum(numbers) / len(numbers)"""

inputs = tokenizer("summarize code: " + code, return_tensors="pt",
                   truncation=True, max_length=512)
# Decoding settings (beam search etc.) are loaded from generation_config.json.
output = model.generate(**inputs)
print(tokenizer.decode(output[0], skip_special_tokens=True))
```

## Evaluation

Scores on 5,000 held-out test functions, never seen during training or model
selection. Validation scores are from the in-training evaluation subset.

| Metric | Test | Validation |
|---|---|---|
| BLEU (sacreBLEU, corpus) | 3.92 | 4.69 |
| Smoothed BLEU-4 (sentence avg.) | 6.65 | 7.21 |
| ROUGE-1 | 36.03 | 37.42 |
| ROUGE-2 | 12.61 | 14.14 |
| ROUGE-L | 32.97 | 34.02 |
| Semantic similarity (MiniLM cosine) | 54.02 | – |
| Avg. generated length (words) | 6.17 | 5.92 |
| Avg. reference length (words) | 10.02 | 9.90 |

BLEU/ROUGE reward lexical overlap with a single reference docstring, so a correct
summary phrased differently scores low. Read them alongside the examples below.
CodeBLEU is not reported: it scores generated *code*, while this model generates English.

## Examples from the test split

```python
def validate_flavor_data(self, expected, actual):

        self.log.debug('Validating flavor data...')
        self.log.debug('actual: {}'.format(repr(actual)))
        act = [a.name for a in actual]
        return self._validate_list_data(expected, act)
```

- **Reference:** Validate flavor data.
- **Generated:** Validate flavor data.

```python
def check(text):
    err = "hedging.misc"
    msg = "Hedging. Just say it."

    narcissism = [
        "I would argue that",
        ", so to speak",
        "to a certain degree",
    ]

    return existence_check(text, narcissism, err, msg)
```

- **Reference:** Suggest the preferred forms.
- **Generated:** Check if hedging is valid.

```python
def on_source_directory_chooser_clicked(self):

        title = self.tr('Set the source directory for script and scenario')
        self.choose_directory(self.source_directory, title)
```

- **Reference:** Autoconnect slot activated when tbSourceDir is clicked.
- **Generated:** Sets the source directory for script and scenario.

## Training data

[`sentence-transformers/codesearchnet`](https://huggingface.co/datasets/sentence-transformers/codesearchnet) (`pair` config),
`code` → `comment` pairs. The dataset mixes about six languages without a label, so
Python functions were detected by parsing with `ast`. Leading docstrings were stripped
from the code (otherwise the target leaks into the input), summaries were cut to their
leading prose, and broken, non-English and boilerplate rows were dropped.

| Split | Examples |
|---|---|
| train | 20,000 |
| validation | 5,000 |
| test | 5,000 |

## Training procedure

| Hyper-parameter | Value |
|---|---|
| Learning rate | 0.0003 |
| Scheduler / warmup | linear / 0.03 |
| Optimizer | adamw_torch |
| Effective batch size | 32 (per-device 8 × accumulation 4) |
| Epochs | 3.0 |
| Weight decay | 0.01 |
| Max source / target length | 512 / 64 tokens |
| Precision | bf16 |
| Gradient checkpointing | True |
| Seed | 42 |
| Training time | 50.49 min |
| Peak GPU memory | 4.17 GB |

Best checkpoint selected on validation ROUGE-L with early stopping.

### Generation

`num_beams=4`, `max_length=64`, `min_length=4`, `length_penalty=1.0`, `no_repeat_ngram_size=3`, `early_stopping=True`, `do_sample=False`

## Limitations

- Trained on Python only; other languages are out of distribution.
- Inputs longer than 512 tokens are truncated, so the end of long functions is not seen.
- Summaries tend to be shorter and more generic than human-written docstrings.
- Docstrings in CodeSearchNet are noisy; the model inherits their style and errors.