PedramR commited on
Commit
c9216d2
·
verified ·
1 Parent(s): 61f22a8

checkpoint-1000

Browse files
qwen3.5-2b/checkpoint-1000/README.md ADDED
@@ -0,0 +1,207 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: Qwen/Qwen3.5-2B
3
+ library_name: peft
4
+ pipeline_tag: text-generation
5
+ tags:
6
+ - base_model:adapter:Qwen/Qwen3.5-2B
7
+ - lora
8
+ - transformers
9
+ ---
10
+
11
+ # Model Card for Model ID
12
+
13
+ <!-- Provide a quick summary of what the model is/does. -->
14
+
15
+
16
+
17
+ ## Model Details
18
+
19
+ ### Model Description
20
+
21
+ <!-- Provide a longer summary of what this model is. -->
22
+
23
+
24
+
25
+ - **Developed by:** [More Information Needed]
26
+ - **Funded by [optional]:** [More Information Needed]
27
+ - **Shared by [optional]:** [More Information Needed]
28
+ - **Model type:** [More Information Needed]
29
+ - **Language(s) (NLP):** [More Information Needed]
30
+ - **License:** [More Information Needed]
31
+ - **Finetuned from model [optional]:** [More Information Needed]
32
+
33
+ ### Model Sources [optional]
34
+
35
+ <!-- Provide the basic links for the model. -->
36
+
37
+ - **Repository:** [More Information Needed]
38
+ - **Paper [optional]:** [More Information Needed]
39
+ - **Demo [optional]:** [More Information Needed]
40
+
41
+ ## Uses
42
+
43
+ <!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
44
+
45
+ ### Direct Use
46
+
47
+ <!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
48
+
49
+ [More Information Needed]
50
+
51
+ ### Downstream Use [optional]
52
+
53
+ <!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
54
+
55
+ [More Information Needed]
56
+
57
+ ### Out-of-Scope Use
58
+
59
+ <!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
60
+
61
+ [More Information Needed]
62
+
63
+ ## Bias, Risks, and Limitations
64
+
65
+ <!-- This section is meant to convey both technical and sociotechnical limitations. -->
66
+
67
+ [More Information Needed]
68
+
69
+ ### Recommendations
70
+
71
+ <!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
72
+
73
+ Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
74
+
75
+ ## How to Get Started with the Model
76
+
77
+ Use the code below to get started with the model.
78
+
79
+ [More Information Needed]
80
+
81
+ ## Training Details
82
+
83
+ ### Training Data
84
+
85
+ <!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
86
+
87
+ [More Information Needed]
88
+
89
+ ### Training Procedure
90
+
91
+ <!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
92
+
93
+ #### Preprocessing [optional]
94
+
95
+ [More Information Needed]
96
+
97
+
98
+ #### Training Hyperparameters
99
+
100
+ - **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
101
+
102
+ #### Speeds, Sizes, Times [optional]
103
+
104
+ <!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
105
+
106
+ [More Information Needed]
107
+
108
+ ## Evaluation
109
+
110
+ <!-- This section describes the evaluation protocols and provides the results. -->
111
+
112
+ ### Testing Data, Factors & Metrics
113
+
114
+ #### Testing Data
115
+
116
+ <!-- This should link to a Dataset Card if possible. -->
117
+
118
+ [More Information Needed]
119
+
120
+ #### Factors
121
+
122
+ <!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
123
+
124
+ [More Information Needed]
125
+
126
+ #### Metrics
127
+
128
+ <!-- These are the evaluation metrics being used, ideally with a description of why. -->
129
+
130
+ [More Information Needed]
131
+
132
+ ### Results
133
+
134
+ [More Information Needed]
135
+
136
+ #### Summary
137
+
138
+
139
+
140
+ ## Model Examination [optional]
141
+
142
+ <!-- Relevant interpretability work for the model goes here -->
143
+
144
+ [More Information Needed]
145
+
146
+ ## Environmental Impact
147
+
148
+ <!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
149
+
150
+ Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
151
+
152
+ - **Hardware Type:** [More Information Needed]
153
+ - **Hours used:** [More Information Needed]
154
+ - **Cloud Provider:** [More Information Needed]
155
+ - **Compute Region:** [More Information Needed]
156
+ - **Carbon Emitted:** [More Information Needed]
157
+
158
+ ## Technical Specifications [optional]
159
+
160
+ ### Model Architecture and Objective
161
+
162
+ [More Information Needed]
163
+
164
+ ### Compute Infrastructure
165
+
166
+ [More Information Needed]
167
+
168
+ #### Hardware
169
+
170
+ [More Information Needed]
171
+
172
+ #### Software
173
+
174
+ [More Information Needed]
175
+
176
+ ## Citation [optional]
177
+
178
+ <!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
179
+
180
+ **BibTeX:**
181
+
182
+ [More Information Needed]
183
+
184
+ **APA:**
185
+
186
+ [More Information Needed]
187
+
188
+ ## Glossary [optional]
189
+
190
+ <!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
191
+
192
+ [More Information Needed]
193
+
194
+ ## More Information [optional]
195
+
196
+ [More Information Needed]
197
+
198
+ ## Model Card Authors [optional]
199
+
200
+ [More Information Needed]
201
+
202
+ ## Model Card Contact
203
+
204
+ [More Information Needed]
205
+ ### Framework versions
206
+
207
+ - PEFT 0.21.0
qwen3.5-2b/checkpoint-1000/adapter_config.json ADDED
@@ -0,0 +1,56 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "alora_invocation_tokens": null,
3
+ "alpha_pattern": {},
4
+ "arrow_config": null,
5
+ "auto_mapping": null,
6
+ "base_model_name_or_path": "Qwen/Qwen3.5-2B",
7
+ "bias": "none",
8
+ "corda_config": null,
9
+ "ensure_weight_tying": false,
10
+ "eva_config": null,
11
+ "exclude_modules": null,
12
+ "fan_in_fan_out": false,
13
+ "inference_mode": true,
14
+ "init_lora_weights": true,
15
+ "kasa_config": null,
16
+ "layer_replication": null,
17
+ "layers_pattern": null,
18
+ "layers_to_transform": null,
19
+ "loftq_config": {},
20
+ "lora_alpha": 64,
21
+ "lora_bias": false,
22
+ "lora_dropout": 0.05,
23
+ "lora_ga_config": null,
24
+ "megatron_config": null,
25
+ "megatron_core": "megatron.core",
26
+ "modules_to_save": null,
27
+ "monteclora_config": null,
28
+ "peft_type": "LORA",
29
+ "peft_version": "0.21.0",
30
+ "qalora_group_size": 16,
31
+ "r": 32,
32
+ "rank_pattern": {},
33
+ "revision": null,
34
+ "target_modules": [
35
+ "up_proj",
36
+ "down_proj",
37
+ "gate_proj",
38
+ "out_proj",
39
+ "in_proj_a",
40
+ "in_proj_qkv",
41
+ "in_proj_z",
42
+ "o_proj",
43
+ "in_proj_b",
44
+ "v_proj",
45
+ "k_proj",
46
+ "q_proj"
47
+ ],
48
+ "target_parameters": null,
49
+ "task_type": "CAUSAL_LM",
50
+ "trainable_token_indices": null,
51
+ "use_bdlora": null,
52
+ "use_dora": false,
53
+ "use_qalora": false,
54
+ "use_rslora": false,
55
+ "velora_config": null
56
+ }
qwen3.5-2b/checkpoint-1000/adapter_model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a0e20a3df23bdcd723f671ce34b80795b783e0d43b77af5ea86cd54632057f9f
3
+ size 134604152
qwen3.5-2b/checkpoint-1000/optimizer.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:eb40f5b419fb45a01f27564d943532819216044683eff73c8b9a94ed898c8b43
3
+ size 269426431
qwen3.5-2b/checkpoint-1000/rng_state.pth ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a1038a49bd92fa68e4a6e653212221a93e9b68affa3208f57f91ea09f7fe28e2
3
+ size 14645
qwen3.5-2b/checkpoint-1000/scheduler.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6fd37ccb5616025d03cf39ffeb34cf2981e7ff5ce5b5ddce7468fa267198b72b
3
+ size 1465
qwen3.5-2b/checkpoint-1000/trainer_state.json ADDED
@@ -0,0 +1,758 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "best_global_step": null,
3
+ "best_metric": null,
4
+ "best_model_checkpoint": null,
5
+ "epoch": 0.5278785879247773,
6
+ "eval_steps": 500,
7
+ "global_step": 1000,
8
+ "is_hyper_param_search": false,
9
+ "is_local_process_zero": true,
10
+ "is_world_process_zero": true,
11
+ "log_history": [
12
+ {
13
+ "epoch": 0.005278785879247773,
14
+ "grad_norm": 1.305302381515503,
15
+ "learning_rate": 1.8947368421052634e-05,
16
+ "loss": 0.8562479972839355,
17
+ "step": 10
18
+ },
19
+ {
20
+ "epoch": 0.010557571758495546,
21
+ "grad_norm": 0.9230896830558777,
22
+ "learning_rate": 4e-05,
23
+ "loss": 0.7965596675872803,
24
+ "step": 20
25
+ },
26
+ {
27
+ "epoch": 0.01583635763774332,
28
+ "grad_norm": 0.7995359301567078,
29
+ "learning_rate": 6.105263157894737e-05,
30
+ "loss": 0.7992881298065185,
31
+ "step": 30
32
+ },
33
+ {
34
+ "epoch": 0.02111514351699109,
35
+ "grad_norm": 0.7340251207351685,
36
+ "learning_rate": 8.210526315789474e-05,
37
+ "loss": 0.6521625995635987,
38
+ "step": 40
39
+ },
40
+ {
41
+ "epoch": 0.026393929396238865,
42
+ "grad_norm": 1.0072991847991943,
43
+ "learning_rate": 0.00010315789473684211,
44
+ "loss": 0.6839987277984619,
45
+ "step": 50
46
+ },
47
+ {
48
+ "epoch": 0.03167271527548664,
49
+ "grad_norm": 0.8970349431037903,
50
+ "learning_rate": 0.00012421052631578949,
51
+ "loss": 0.6341384887695313,
52
+ "step": 60
53
+ },
54
+ {
55
+ "epoch": 0.03695150115473441,
56
+ "grad_norm": 0.883733332157135,
57
+ "learning_rate": 0.00014526315789473686,
58
+ "loss": 0.6603663444519043,
59
+ "step": 70
60
+ },
61
+ {
62
+ "epoch": 0.04223028703398218,
63
+ "grad_norm": 0.8134523630142212,
64
+ "learning_rate": 0.00016631578947368423,
65
+ "loss": 0.6564276695251465,
66
+ "step": 80
67
+ },
68
+ {
69
+ "epoch": 0.047509072913229956,
70
+ "grad_norm": 0.8890961408615112,
71
+ "learning_rate": 0.0001873684210526316,
72
+ "loss": 0.6585474967956543,
73
+ "step": 90
74
+ },
75
+ {
76
+ "epoch": 0.05278785879247773,
77
+ "grad_norm": 0.7649273872375488,
78
+ "learning_rate": 0.00019955555555555558,
79
+ "loss": 0.6272543430328369,
80
+ "step": 100
81
+ },
82
+ {
83
+ "epoch": 0.0580666446717255,
84
+ "grad_norm": 0.7554565072059631,
85
+ "learning_rate": 0.00019844444444444445,
86
+ "loss": 0.5757376194000244,
87
+ "step": 110
88
+ },
89
+ {
90
+ "epoch": 0.06334543055097328,
91
+ "grad_norm": 0.7954307794570923,
92
+ "learning_rate": 0.00019733333333333335,
93
+ "loss": 0.641472053527832,
94
+ "step": 120
95
+ },
96
+ {
97
+ "epoch": 0.06862421643022105,
98
+ "grad_norm": 0.672359824180603,
99
+ "learning_rate": 0.00019622222222222225,
100
+ "loss": 0.563064432144165,
101
+ "step": 130
102
+ },
103
+ {
104
+ "epoch": 0.07390300230946882,
105
+ "grad_norm": 0.6992788314819336,
106
+ "learning_rate": 0.0001951111111111111,
107
+ "loss": 0.600986099243164,
108
+ "step": 140
109
+ },
110
+ {
111
+ "epoch": 0.0791817881887166,
112
+ "grad_norm": 0.6328181028366089,
113
+ "learning_rate": 0.000194,
114
+ "loss": 0.5295759201049804,
115
+ "step": 150
116
+ },
117
+ {
118
+ "epoch": 0.08446057406796437,
119
+ "grad_norm": 0.6692398190498352,
120
+ "learning_rate": 0.0001928888888888889,
121
+ "loss": 0.6027700901031494,
122
+ "step": 160
123
+ },
124
+ {
125
+ "epoch": 0.08973935994721215,
126
+ "grad_norm": 0.7081276774406433,
127
+ "learning_rate": 0.0001917777777777778,
128
+ "loss": 0.5615420341491699,
129
+ "step": 170
130
+ },
131
+ {
132
+ "epoch": 0.09501814582645991,
133
+ "grad_norm": 0.5484495759010315,
134
+ "learning_rate": 0.00019066666666666668,
135
+ "loss": 0.5883731842041016,
136
+ "step": 180
137
+ },
138
+ {
139
+ "epoch": 0.10029693170570769,
140
+ "grad_norm": 0.7612242102622986,
141
+ "learning_rate": 0.00018955555555555558,
142
+ "loss": 0.5624749183654785,
143
+ "step": 190
144
+ },
145
+ {
146
+ "epoch": 0.10557571758495546,
147
+ "grad_norm": 0.678548276424408,
148
+ "learning_rate": 0.00018844444444444445,
149
+ "loss": 0.5546360969543457,
150
+ "step": 200
151
+ },
152
+ {
153
+ "epoch": 0.11085450346420324,
154
+ "grad_norm": 0.5756900906562805,
155
+ "learning_rate": 0.00018733333333333335,
156
+ "loss": 0.5600387573242187,
157
+ "step": 210
158
+ },
159
+ {
160
+ "epoch": 0.116133289343451,
161
+ "grad_norm": 0.5475565791130066,
162
+ "learning_rate": 0.00018622222222222223,
163
+ "loss": 0.5820147037506104,
164
+ "step": 220
165
+ },
166
+ {
167
+ "epoch": 0.12141207522269878,
168
+ "grad_norm": 0.6964915990829468,
169
+ "learning_rate": 0.00018511111111111113,
170
+ "loss": 0.5699102401733398,
171
+ "step": 230
172
+ },
173
+ {
174
+ "epoch": 0.12669086110194655,
175
+ "grad_norm": 0.6508777737617493,
176
+ "learning_rate": 0.00018400000000000003,
177
+ "loss": 0.5863839626312256,
178
+ "step": 240
179
+ },
180
+ {
181
+ "epoch": 0.13196964698119432,
182
+ "grad_norm": 0.6375535726547241,
183
+ "learning_rate": 0.00018288888888888887,
184
+ "loss": 0.5772487640380859,
185
+ "step": 250
186
+ },
187
+ {
188
+ "epoch": 0.1372484328604421,
189
+ "grad_norm": 0.5595227479934692,
190
+ "learning_rate": 0.00018177777777777778,
191
+ "loss": 0.603016710281372,
192
+ "step": 260
193
+ },
194
+ {
195
+ "epoch": 0.14252721873968988,
196
+ "grad_norm": 0.5612871646881104,
197
+ "learning_rate": 0.00018066666666666668,
198
+ "loss": 0.5080572605133057,
199
+ "step": 270
200
+ },
201
+ {
202
+ "epoch": 0.14780600461893764,
203
+ "grad_norm": 0.762787938117981,
204
+ "learning_rate": 0.00017955555555555558,
205
+ "loss": 0.6409445762634277,
206
+ "step": 280
207
+ },
208
+ {
209
+ "epoch": 0.1530847904981854,
210
+ "grad_norm": 0.6885519623756409,
211
+ "learning_rate": 0.00017844444444444445,
212
+ "loss": 0.6149398803710937,
213
+ "step": 290
214
+ },
215
+ {
216
+ "epoch": 0.1583635763774332,
217
+ "grad_norm": 0.6986534595489502,
218
+ "learning_rate": 0.00017733333333333335,
219
+ "loss": 0.5259696960449218,
220
+ "step": 300
221
+ },
222
+ {
223
+ "epoch": 0.16364236225668097,
224
+ "grad_norm": 0.601087212562561,
225
+ "learning_rate": 0.00017622222222222223,
226
+ "loss": 0.6464223384857177,
227
+ "step": 310
228
+ },
229
+ {
230
+ "epoch": 0.16892114813592873,
231
+ "grad_norm": 0.4432871639728546,
232
+ "learning_rate": 0.00017511111111111113,
233
+ "loss": 0.509718942642212,
234
+ "step": 320
235
+ },
236
+ {
237
+ "epoch": 0.1741999340151765,
238
+ "grad_norm": 0.5295445322990417,
239
+ "learning_rate": 0.000174,
240
+ "loss": 0.5282030582427979,
241
+ "step": 330
242
+ },
243
+ {
244
+ "epoch": 0.1794787198944243,
245
+ "grad_norm": 0.695189893245697,
246
+ "learning_rate": 0.0001728888888888889,
247
+ "loss": 0.5390444278717041,
248
+ "step": 340
249
+ },
250
+ {
251
+ "epoch": 0.18475750577367206,
252
+ "grad_norm": 0.6871212124824524,
253
+ "learning_rate": 0.0001717777777777778,
254
+ "loss": 0.6001670360565186,
255
+ "step": 350
256
+ },
257
+ {
258
+ "epoch": 0.19003629165291983,
259
+ "grad_norm": 0.5886061787605286,
260
+ "learning_rate": 0.00017066666666666668,
261
+ "loss": 0.5532281875610352,
262
+ "step": 360
263
+ },
264
+ {
265
+ "epoch": 0.1953150775321676,
266
+ "grad_norm": 0.4359692931175232,
267
+ "learning_rate": 0.00016955555555555555,
268
+ "loss": 0.5710611343383789,
269
+ "step": 370
270
+ },
271
+ {
272
+ "epoch": 0.20059386341141539,
273
+ "grad_norm": 0.5859516859054565,
274
+ "learning_rate": 0.00016844444444444445,
275
+ "loss": 0.527172040939331,
276
+ "step": 380
277
+ },
278
+ {
279
+ "epoch": 0.20587264929066315,
280
+ "grad_norm": 0.595167875289917,
281
+ "learning_rate": 0.00016733333333333335,
282
+ "loss": 0.5060821056365967,
283
+ "step": 390
284
+ },
285
+ {
286
+ "epoch": 0.21115143516991092,
287
+ "grad_norm": 0.5597789883613586,
288
+ "learning_rate": 0.00016622222222222223,
289
+ "loss": 0.5472925186157227,
290
+ "step": 400
291
+ },
292
+ {
293
+ "epoch": 0.21643022104915868,
294
+ "grad_norm": 0.5002998113632202,
295
+ "learning_rate": 0.00016511111111111113,
296
+ "loss": 0.5917861461639404,
297
+ "step": 410
298
+ },
299
+ {
300
+ "epoch": 0.22170900692840648,
301
+ "grad_norm": 0.6672388911247253,
302
+ "learning_rate": 0.000164,
303
+ "loss": 0.5691781044006348,
304
+ "step": 420
305
+ },
306
+ {
307
+ "epoch": 0.22698779280765424,
308
+ "grad_norm": 0.580162763595581,
309
+ "learning_rate": 0.0001628888888888889,
310
+ "loss": 0.5123159408569335,
311
+ "step": 430
312
+ },
313
+ {
314
+ "epoch": 0.232266578686902,
315
+ "grad_norm": 0.5568054914474487,
316
+ "learning_rate": 0.00016177777777777778,
317
+ "loss": 0.5455626010894775,
318
+ "step": 440
319
+ },
320
+ {
321
+ "epoch": 0.23754536456614977,
322
+ "grad_norm": 0.6651593446731567,
323
+ "learning_rate": 0.00016066666666666668,
324
+ "loss": 0.5399269104003906,
325
+ "step": 450
326
+ },
327
+ {
328
+ "epoch": 0.24282415044539757,
329
+ "grad_norm": 0.5001797080039978,
330
+ "learning_rate": 0.00015955555555555558,
331
+ "loss": 0.5277121543884278,
332
+ "step": 460
333
+ },
334
+ {
335
+ "epoch": 0.24810293632464533,
336
+ "grad_norm": 0.5514921545982361,
337
+ "learning_rate": 0.00015844444444444445,
338
+ "loss": 0.5342081546783447,
339
+ "step": 470
340
+ },
341
+ {
342
+ "epoch": 0.2533817222038931,
343
+ "grad_norm": 0.5599192380905151,
344
+ "learning_rate": 0.00015733333333333333,
345
+ "loss": 0.5150089263916016,
346
+ "step": 480
347
+ },
348
+ {
349
+ "epoch": 0.2586605080831409,
350
+ "grad_norm": 0.563194215297699,
351
+ "learning_rate": 0.00015622222222222223,
352
+ "loss": 0.5231950759887696,
353
+ "step": 490
354
+ },
355
+ {
356
+ "epoch": 0.26393929396238863,
357
+ "grad_norm": 0.7210673689842224,
358
+ "learning_rate": 0.00015511111111111113,
359
+ "loss": 0.5708573818206787,
360
+ "step": 500
361
+ },
362
+ {
363
+ "epoch": 0.26393929396238863,
364
+ "eval_loss": 0.534747838973999,
365
+ "eval_runtime": 459.3797,
366
+ "eval_samples_per_second": 24.383,
367
+ "eval_steps_per_second": 6.097,
368
+ "step": 500
369
+ },
370
+ {
371
+ "epoch": 0.26393929396238863,
372
+ "step": 500,
373
+ "test_best_wer": 0.21571711227233525,
374
+ "test_exact_match": 0.0,
375
+ "test_hallucination_rate": 0.09375,
376
+ "test_wer": 0.4919080716351276
377
+ },
378
+ {
379
+ "epoch": 0.2692180798416364,
380
+ "grad_norm": 0.5391792058944702,
381
+ "learning_rate": 0.000154,
382
+ "loss": 0.5354532241821289,
383
+ "step": 510
384
+ },
385
+ {
386
+ "epoch": 0.2744968657208842,
387
+ "grad_norm": 0.5327297449111938,
388
+ "learning_rate": 0.0001528888888888889,
389
+ "loss": 0.5177035331726074,
390
+ "step": 520
391
+ },
392
+ {
393
+ "epoch": 0.27977565160013196,
394
+ "grad_norm": 0.5588216781616211,
395
+ "learning_rate": 0.00015177777777777778,
396
+ "loss": 0.5256490707397461,
397
+ "step": 530
398
+ },
399
+ {
400
+ "epoch": 0.28505443747937975,
401
+ "grad_norm": 0.6024046540260315,
402
+ "learning_rate": 0.00015066666666666668,
403
+ "loss": 0.5359174728393554,
404
+ "step": 540
405
+ },
406
+ {
407
+ "epoch": 0.2903332233586275,
408
+ "grad_norm": 0.5452874898910522,
409
+ "learning_rate": 0.00014955555555555555,
410
+ "loss": 0.5050687313079834,
411
+ "step": 550
412
+ },
413
+ {
414
+ "epoch": 0.2956120092378753,
415
+ "grad_norm": 0.5716919898986816,
416
+ "learning_rate": 0.00014844444444444445,
417
+ "loss": 0.5043536186218261,
418
+ "step": 560
419
+ },
420
+ {
421
+ "epoch": 0.3008907951171231,
422
+ "grad_norm": 0.5621301531791687,
423
+ "learning_rate": 0.00014733333333333335,
424
+ "loss": 0.5681197643280029,
425
+ "step": 570
426
+ },
427
+ {
428
+ "epoch": 0.3061695809963708,
429
+ "grad_norm": 0.49813947081565857,
430
+ "learning_rate": 0.00014622222222222223,
431
+ "loss": 0.5073559761047364,
432
+ "step": 580
433
+ },
434
+ {
435
+ "epoch": 0.3114483668756186,
436
+ "grad_norm": 0.5197216272354126,
437
+ "learning_rate": 0.0001451111111111111,
438
+ "loss": 0.510869026184082,
439
+ "step": 590
440
+ },
441
+ {
442
+ "epoch": 0.3167271527548664,
443
+ "grad_norm": 0.5098935961723328,
444
+ "learning_rate": 0.000144,
445
+ "loss": 0.5177080154418945,
446
+ "step": 600
447
+ },
448
+ {
449
+ "epoch": 0.32200593863411414,
450
+ "grad_norm": 0.5004557967185974,
451
+ "learning_rate": 0.0001428888888888889,
452
+ "loss": 0.5292775154113769,
453
+ "step": 610
454
+ },
455
+ {
456
+ "epoch": 0.32728472451336194,
457
+ "grad_norm": 0.6179164052009583,
458
+ "learning_rate": 0.00014177777777777778,
459
+ "loss": 0.532755994796753,
460
+ "step": 620
461
+ },
462
+ {
463
+ "epoch": 0.3325635103926097,
464
+ "grad_norm": 0.4971306025981903,
465
+ "learning_rate": 0.00014066666666666668,
466
+ "loss": 0.5370781421661377,
467
+ "step": 630
468
+ },
469
+ {
470
+ "epoch": 0.33784229627185747,
471
+ "grad_norm": 0.5876296758651733,
472
+ "learning_rate": 0.00013955555555555558,
473
+ "loss": 0.5408426284790039,
474
+ "step": 640
475
+ },
476
+ {
477
+ "epoch": 0.34312108215110526,
478
+ "grad_norm": 0.4632679522037506,
479
+ "learning_rate": 0.00013844444444444445,
480
+ "loss": 0.475573205947876,
481
+ "step": 650
482
+ },
483
+ {
484
+ "epoch": 0.348399868030353,
485
+ "grad_norm": 0.6051201224327087,
486
+ "learning_rate": 0.00013733333333333333,
487
+ "loss": 0.5176257133483887,
488
+ "step": 660
489
+ },
490
+ {
491
+ "epoch": 0.3536786539096008,
492
+ "grad_norm": 0.5086127519607544,
493
+ "learning_rate": 0.00013622222222222223,
494
+ "loss": 0.543503713607788,
495
+ "step": 670
496
+ },
497
+ {
498
+ "epoch": 0.3589574397888486,
499
+ "grad_norm": 0.45412692427635193,
500
+ "learning_rate": 0.00013511111111111113,
501
+ "loss": 0.4648271560668945,
502
+ "step": 680
503
+ },
504
+ {
505
+ "epoch": 0.3642362256680963,
506
+ "grad_norm": 0.5843290686607361,
507
+ "learning_rate": 0.000134,
508
+ "loss": 0.4611194133758545,
509
+ "step": 690
510
+ },
511
+ {
512
+ "epoch": 0.3695150115473441,
513
+ "grad_norm": 0.5016586184501648,
514
+ "learning_rate": 0.00013288888888888888,
515
+ "loss": 0.5229296684265137,
516
+ "step": 700
517
+ },
518
+ {
519
+ "epoch": 0.37479379742659186,
520
+ "grad_norm": 0.594906747341156,
521
+ "learning_rate": 0.00013177777777777778,
522
+ "loss": 0.4897459506988525,
523
+ "step": 710
524
+ },
525
+ {
526
+ "epoch": 0.38007258330583965,
527
+ "grad_norm": 0.6894946098327637,
528
+ "learning_rate": 0.00013066666666666668,
529
+ "loss": 0.5067539691925049,
530
+ "step": 720
531
+ },
532
+ {
533
+ "epoch": 0.38535136918508744,
534
+ "grad_norm": 0.5899446606636047,
535
+ "learning_rate": 0.00012955555555555555,
536
+ "loss": 0.572650671005249,
537
+ "step": 730
538
+ },
539
+ {
540
+ "epoch": 0.3906301550643352,
541
+ "grad_norm": 0.5914052724838257,
542
+ "learning_rate": 0.00012844444444444446,
543
+ "loss": 0.49350743293762206,
544
+ "step": 740
545
+ },
546
+ {
547
+ "epoch": 0.395908940943583,
548
+ "grad_norm": 0.5701265931129456,
549
+ "learning_rate": 0.00012733333333333336,
550
+ "loss": 0.4886340141296387,
551
+ "step": 750
552
+ },
553
+ {
554
+ "epoch": 0.40118772682283077,
555
+ "grad_norm": 0.5554320216178894,
556
+ "learning_rate": 0.00012622222222222223,
557
+ "loss": 0.5152291774749755,
558
+ "step": 760
559
+ },
560
+ {
561
+ "epoch": 0.4064665127020785,
562
+ "grad_norm": 0.542721688747406,
563
+ "learning_rate": 0.0001251111111111111,
564
+ "loss": 0.5169697284698487,
565
+ "step": 770
566
+ },
567
+ {
568
+ "epoch": 0.4117452985813263,
569
+ "grad_norm": 0.5239601731300354,
570
+ "learning_rate": 0.000124,
571
+ "loss": 0.5226495265960693,
572
+ "step": 780
573
+ },
574
+ {
575
+ "epoch": 0.41702408446057404,
576
+ "grad_norm": 0.5628821849822998,
577
+ "learning_rate": 0.0001228888888888889,
578
+ "loss": 0.4952070236206055,
579
+ "step": 790
580
+ },
581
+ {
582
+ "epoch": 0.42230287033982183,
583
+ "grad_norm": 0.53690505027771,
584
+ "learning_rate": 0.0001217777777777778,
585
+ "loss": 0.5176570892333985,
586
+ "step": 800
587
+ },
588
+ {
589
+ "epoch": 0.42758165621906963,
590
+ "grad_norm": 0.47472912073135376,
591
+ "learning_rate": 0.00012066666666666668,
592
+ "loss": 0.48554182052612305,
593
+ "step": 810
594
+ },
595
+ {
596
+ "epoch": 0.43286044209831737,
597
+ "grad_norm": 0.5943832397460938,
598
+ "learning_rate": 0.00011955555555555556,
599
+ "loss": 0.5519014835357666,
600
+ "step": 820
601
+ },
602
+ {
603
+ "epoch": 0.43813922797756516,
604
+ "grad_norm": 0.5702099800109863,
605
+ "learning_rate": 0.00011844444444444444,
606
+ "loss": 0.5203097343444825,
607
+ "step": 830
608
+ },
609
+ {
610
+ "epoch": 0.44341801385681295,
611
+ "grad_norm": 0.5570870637893677,
612
+ "learning_rate": 0.00011733333333333334,
613
+ "loss": 0.5008646488189697,
614
+ "step": 840
615
+ },
616
+ {
617
+ "epoch": 0.4486967997360607,
618
+ "grad_norm": 0.6742717623710632,
619
+ "learning_rate": 0.00011622222222222223,
620
+ "loss": 0.4884360313415527,
621
+ "step": 850
622
+ },
623
+ {
624
+ "epoch": 0.4539755856153085,
625
+ "grad_norm": 0.5782440900802612,
626
+ "learning_rate": 0.00011511111111111112,
627
+ "loss": 0.497051477432251,
628
+ "step": 860
629
+ },
630
+ {
631
+ "epoch": 0.4592543714945563,
632
+ "grad_norm": 0.6157666444778442,
633
+ "learning_rate": 0.00011399999999999999,
634
+ "loss": 0.5011940956115722,
635
+ "step": 870
636
+ },
637
+ {
638
+ "epoch": 0.464533157373804,
639
+ "grad_norm": 0.4744451344013214,
640
+ "learning_rate": 0.0001128888888888889,
641
+ "loss": 0.4997218132019043,
642
+ "step": 880
643
+ },
644
+ {
645
+ "epoch": 0.4698119432530518,
646
+ "grad_norm": 0.5648385286331177,
647
+ "learning_rate": 0.00011177777777777778,
648
+ "loss": 0.5077646255493165,
649
+ "step": 890
650
+ },
651
+ {
652
+ "epoch": 0.47509072913229955,
653
+ "grad_norm": 0.5268929600715637,
654
+ "learning_rate": 0.00011066666666666667,
655
+ "loss": 0.4928537368774414,
656
+ "step": 900
657
+ },
658
+ {
659
+ "epoch": 0.48036951501154734,
660
+ "grad_norm": 0.5827873945236206,
661
+ "learning_rate": 0.00010955555555555557,
662
+ "loss": 0.5176413536071778,
663
+ "step": 910
664
+ },
665
+ {
666
+ "epoch": 0.48564830089079514,
667
+ "grad_norm": 0.5056174993515015,
668
+ "learning_rate": 0.00010844444444444446,
669
+ "loss": 0.5180400848388672,
670
+ "step": 920
671
+ },
672
+ {
673
+ "epoch": 0.4909270867700429,
674
+ "grad_norm": 0.47383710741996765,
675
+ "learning_rate": 0.00010733333333333333,
676
+ "loss": 0.4904231071472168,
677
+ "step": 930
678
+ },
679
+ {
680
+ "epoch": 0.49620587264929067,
681
+ "grad_norm": 0.5303201079368591,
682
+ "learning_rate": 0.00010622222222222222,
683
+ "loss": 0.4838510036468506,
684
+ "step": 940
685
+ },
686
+ {
687
+ "epoch": 0.5014846585285384,
688
+ "grad_norm": 0.6136945486068726,
689
+ "learning_rate": 0.00010511111111111112,
690
+ "loss": 0.4908186912536621,
691
+ "step": 950
692
+ },
693
+ {
694
+ "epoch": 0.5067634444077862,
695
+ "grad_norm": 0.6062966585159302,
696
+ "learning_rate": 0.00010400000000000001,
697
+ "loss": 0.5015266895294189,
698
+ "step": 960
699
+ },
700
+ {
701
+ "epoch": 0.512042230287034,
702
+ "grad_norm": 0.5623139142990112,
703
+ "learning_rate": 0.0001028888888888889,
704
+ "loss": 0.5161266803741456,
705
+ "step": 970
706
+ },
707
+ {
708
+ "epoch": 0.5173210161662818,
709
+ "grad_norm": 0.5548786520957947,
710
+ "learning_rate": 0.00010177777777777777,
711
+ "loss": 0.4842812538146973,
712
+ "step": 980
713
+ },
714
+ {
715
+ "epoch": 0.5225998020455296,
716
+ "grad_norm": 0.4839674234390259,
717
+ "learning_rate": 0.00010066666666666667,
718
+ "loss": 0.47931456565856934,
719
+ "step": 990
720
+ },
721
+ {
722
+ "epoch": 0.5278785879247773,
723
+ "grad_norm": 0.6106690764427185,
724
+ "learning_rate": 9.955555555555556e-05,
725
+ "loss": 0.5249813079833985,
726
+ "step": 1000
727
+ },
728
+ {
729
+ "epoch": 0.5278785879247773,
730
+ "eval_loss": 0.4978668689727783,
731
+ "eval_runtime": 459.5591,
732
+ "eval_samples_per_second": 24.373,
733
+ "eval_steps_per_second": 6.095,
734
+ "step": 1000
735
+ }
736
+ ],
737
+ "logging_steps": 10,
738
+ "max_steps": 1895,
739
+ "num_input_tokens_seen": 0,
740
+ "num_train_epochs": 1,
741
+ "save_steps": 500,
742
+ "stateful_callbacks": {
743
+ "TrainerControl": {
744
+ "args": {
745
+ "should_epoch_stop": false,
746
+ "should_evaluate": false,
747
+ "should_log": false,
748
+ "should_save": true,
749
+ "should_training_stop": false
750
+ },
751
+ "attributes": {}
752
+ }
753
+ },
754
+ "total_flos": 1.1957050265403264e+17,
755
+ "train_batch_size": 2,
756
+ "trial_name": null,
757
+ "trial_params": null
758
+ }
qwen3.5-2b/checkpoint-1000/training_args.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:63f5227842d4c60d51e48d582480b968b2a26456e3111e78999104a00e72c35c
3
+ size 5265
qwen3.5-2b/tb/events.out.tfevents.1789942145.fed349279e2c.8211.0 CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:a28bdce6a24aa7629b8f847f17deb349f7296ec63e987d7db3f5fa8e21c9eab0
3
- size 16665
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d8cb59132d4a5f5f32d6bf0c666a029e77792379d7374978e1a5e35c334efd38
3
+ size 27752
qwen3.5-2b/test_eval/metrics.json CHANGED
@@ -1,6 +1,6 @@
1
  {
2
- "wer": 0.4919080716351276,
3
  "exact_match": 0.0,
4
- "hallucination_rate": 0.09375,
5
  "n_examples": 128
6
  }
 
1
  {
2
+ "wer": 0.3787225407133073,
3
  "exact_match": 0.0,
4
+ "hallucination_rate": 0.0546875,
5
  "n_examples": 128
6
  }
qwen3.5-2b/test_eval/predictions.jsonl CHANGED
The diff for this file is too large to render. See raw diff