File size: 6,000 Bytes
c720b53
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
---
pipeline_tag: image-classification
library_name: pytorch
tags:
- fine-grained-image-recognition
- image-classification
- hierarchical-classification
- timm
---

# Hierarchical-Backbones: hierarchical classification across pretrained backbones

These are the checkpoints behind *Revisiting the Backbone, Pretraining and Transferability for Hierarchical Fine-Grained Image Recognition*, CVGIP 2025 (arXiv link to be added). Twenty-one pretrained backbones (ResNet-50, ViT-B, DeiT-B and DeiT-3-B under supervised, self-supervised and vision-language pretraining) fine-tuned for fine-grained recognition with a hierarchical classification head. The head predicts every level of the annotated label hierarchy, coarsest first: three levels on CUB and Aircraft, two on Cars. Code: [arkel23/Hierarchical](https://github.com/arkel23/Hierarchical).

168 checkpoints, one per configuration, each the last epoch of one training run at 448 px.
Each file is a `torch.save` dict with `config` (the full training configuration), `model` (the
state dict of the backbone and head), `accuracy` and `epoch`, with no optimizer state. File names
are the runs' experiment-log names: dataset, cluster ratio for a pseudo-hierarchy, model, `fz` for
a frozen backbone, and the serial. Load them with
[fgir-zoo](https://github.com/arkel23/fgir-zoo). The [collection](https://huggingface.co/collections/ERISLab/hierarchical-fgir-backbones-and-pretraining-cvgip-2025-6abab163857f4e028b3a38a1) is this work's
page on the Hub; the paper joins it once it is on arXiv.

## Layout

One folder per serial.

| Folder | What | Files | Mean accuracy |
|---|---|---|---|
| `serial_21` | Real (annotated) hierarchy, full fine-tuning, CUB/Aircraft/Cars x 21 backbones, 200 epochs, each backbone at its best learning rate from the paper's sweep, seed 1 | 63 | 90.39 |
| `serial_58` | Aircraft x 21 backbones, 50 epochs, best learning rate, with the real hierarchy; matched with `serial_59` | 21 | 91.39 |
| `serial_59` | Aircraft x 21 backbones, 50 epochs, best learning rate, without a hierarchy; matched with `serial_58` | 21 | 90.94 |
| `serial_60` | Real hierarchy with a frozen backbone (only the classification head trained), the same 63 configurations at their best learning rate, 50 or 200 epochs | 63 | 64.90 |

`manifest.csv` lists every file with its dataset, model, hierarchy (`real`, `pseudo` or `none`),
cluster ratio, frozen backbone, serial, seed, epochs, image size, class count, number of levels,
accuracy, SHA-256 and size.

## Load a checkpoint and classify an image

With a hierarchy the model returns a tuple of logits, one per level from coarsest to finest;
without one it returns a single logits tensor. The finest level is the prediction. The evaluation
transform is the training code's: a bicubic resize to 550 x 550, a 448 px center crop and ImageNet
normalization.

```python
import torch
from PIL import Image
from torchvision import transforms
from fgir_zoo import hierarchical

model = hierarchical.create_model('serial_21/cub_hideit_base_patch16_224.fb_in1k_21')
cfg = model.config
tf = transforms.Compose([
    transforms.Resize((cfg.test_resize_size, cfg.test_resize_size),
                      interpolation=transforms.InterpolationMode.BICUBIC),
    transforms.CenterCrop(cfg.image_size),
    transforms.ToTensor(),
    transforms.Normalize((0.485, 0.456, 0.406), (0.229, 0.224, 0.225)),
])
# a CUB-200-2011 test image, class index 50 (051.Horned_Grebe)
x = tf(Image.open('Horned_Grebe_0050_34561.jpg').convert('RGB')).unsqueeze(0)
with torch.no_grad():
    order, family, species = model(x)  # one logits tensor per level, coarsest first
print(species.argmax(-1).item(), species.softmax(-1).max().item())  # 50 0.8240
```

## Accuracy of the released checkpoints

Top-1 accuracy (%) stored in each file: the run's own finest-level accuracy on the dataset's test
split after the last epoch. The papers report the maximum over a learning-rate sweep, so their
tables differ. Per-file accuracy is in `manifest.csv`; the table below covers
`serial_21`.

| Model | aircraft (21) | cars (21) | cub (21) |
|---|---|---|---|
| `hideit3_base_patch16_224.fb_in1k` | 92.08 | 93.40 | 88.63 |
| `hideit3_base_patch16_224.fb_in22k_ft_in1k` | 92.98 | 93.65 | 90.63 |
| `hideit_base_patch16_224.fb_in1k` | 92.11 | 93.06 | 88.82 |
| `hiresnet50.a1_in1k` | 92.92 | 93.14 | 85.78 |
| `hiresnet50.fb_ssl_yfcc100m_ft_in1k` | 93.07 | 94.28 | 85.55 |
| `hiresnet50.fb_swsl_ig1b_ft_in1k` | 93.10 | 94.28 | 85.48 |
| `hiresnet50.gluon_in1k` | 93.01 | 94.23 | 85.24 |
| `hiresnet50.in1k_mocov3` | 93.07 | 93.53 | 84.85 |
| `hiresnet50.in1k_spark` | 91.27 | 92.38 | 78.63 |
| `hiresnet50.in1k_supcon` | 93.34 | 94.24 | 85.38 |
| `hiresnet50.in1k_swav` | 92.65 | 93.86 | 86.02 |
| `hiresnet50.in21k_miil` | 93.01 | 93.74 | 87.85 |
| `hiresnet50.tv2_in1k` | 92.80 | 93.81 | 87.28 |
| `hiresnet50.tv_in1k` | 92.71 | 94.09 | 86.16 |
| `hivit_base_patch16_224.dino` | 88.00 | 91.10 | 86.47 |
| `hivit_base_patch16_224.in1k_mocov3` | 90.19 | 93.07 | 87.73 |
| `hivit_base_patch16_224.mae` | 83.62 | 92.79 | 85.26 |
| `hivit_base_patch16_224.orig_in21k` | 89.32 | 91.73 | 90.97 |
| `hivit_base_patch16_224_miil.in21k` | 90.49 | 92.13 | 90.21 |
| `hivit_base_patch16_clip_224.laion2b` | 86.41 | 92.36 | 80.45 |
| `hivit_base_patch16_siglip_224.v2_webli` | 93.40 | 95.06 | 87.81 |

## Requirements

- `fgir-zoo` (`pip install git+https://github.com/arkel23/fgir-zoo.git`), which pins
  `timm==0.9.12`
- `torch` (checked with 2.5.1)

## Citation

```bibtex
@inproceedings{surya_revisiting_2025,
  title     = {Revisiting the Backbone, Pretraining and Transferability for Hierarchical
               Fine-Grained Image Recognition},
  author    = {Surya, Augusto Christian and Rios, Edwin Arkel and Lai, Bo-Cheng and Hu, Min-Chun},
  booktitle = {Computer Vision, Graphics, and Image Processing (CVGIP)},
  year      = {2025},
  eprint    = {TBA},
  note      = {A. C. Surya and E. A. Rios contributed equally. arXiv ID to be added.}
}
```