Instructions to use ERISLab/Hierarchical-Backbones with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- timm
How to use ERISLab/Hierarchical-Backbones with timm:
import timm model = timm.create_model("hf_hub:ERISLab/Hierarchical-Backbones", pretrained=True) - Notebooks
- Google Colab
- Kaggle
File size: 6,000 Bytes
c720b53 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 | ---
pipeline_tag: image-classification
library_name: pytorch
tags:
- fine-grained-image-recognition
- image-classification
- hierarchical-classification
- timm
---
# Hierarchical-Backbones: hierarchical classification across pretrained backbones
These are the checkpoints behind *Revisiting the Backbone, Pretraining and Transferability for Hierarchical Fine-Grained Image Recognition*, CVGIP 2025 (arXiv link to be added). Twenty-one pretrained backbones (ResNet-50, ViT-B, DeiT-B and DeiT-3-B under supervised, self-supervised and vision-language pretraining) fine-tuned for fine-grained recognition with a hierarchical classification head. The head predicts every level of the annotated label hierarchy, coarsest first: three levels on CUB and Aircraft, two on Cars. Code: [arkel23/Hierarchical](https://github.com/arkel23/Hierarchical).
168 checkpoints, one per configuration, each the last epoch of one training run at 448 px.
Each file is a `torch.save` dict with `config` (the full training configuration), `model` (the
state dict of the backbone and head), `accuracy` and `epoch`, with no optimizer state. File names
are the runs' experiment-log names: dataset, cluster ratio for a pseudo-hierarchy, model, `fz` for
a frozen backbone, and the serial. Load them with
[fgir-zoo](https://github.com/arkel23/fgir-zoo). The [collection](https://huggingface.co/collections/ERISLab/hierarchical-fgir-backbones-and-pretraining-cvgip-2025-6abab163857f4e028b3a38a1) is this work's
page on the Hub; the paper joins it once it is on arXiv.
## Layout
One folder per serial.
| Folder | What | Files | Mean accuracy |
|---|---|---|---|
| `serial_21` | Real (annotated) hierarchy, full fine-tuning, CUB/Aircraft/Cars x 21 backbones, 200 epochs, each backbone at its best learning rate from the paper's sweep, seed 1 | 63 | 90.39 |
| `serial_58` | Aircraft x 21 backbones, 50 epochs, best learning rate, with the real hierarchy; matched with `serial_59` | 21 | 91.39 |
| `serial_59` | Aircraft x 21 backbones, 50 epochs, best learning rate, without a hierarchy; matched with `serial_58` | 21 | 90.94 |
| `serial_60` | Real hierarchy with a frozen backbone (only the classification head trained), the same 63 configurations at their best learning rate, 50 or 200 epochs | 63 | 64.90 |
`manifest.csv` lists every file with its dataset, model, hierarchy (`real`, `pseudo` or `none`),
cluster ratio, frozen backbone, serial, seed, epochs, image size, class count, number of levels,
accuracy, SHA-256 and size.
## Load a checkpoint and classify an image
With a hierarchy the model returns a tuple of logits, one per level from coarsest to finest;
without one it returns a single logits tensor. The finest level is the prediction. The evaluation
transform is the training code's: a bicubic resize to 550 x 550, a 448 px center crop and ImageNet
normalization.
```python
import torch
from PIL import Image
from torchvision import transforms
from fgir_zoo import hierarchical
model = hierarchical.create_model('serial_21/cub_hideit_base_patch16_224.fb_in1k_21')
cfg = model.config
tf = transforms.Compose([
transforms.Resize((cfg.test_resize_size, cfg.test_resize_size),
interpolation=transforms.InterpolationMode.BICUBIC),
transforms.CenterCrop(cfg.image_size),
transforms.ToTensor(),
transforms.Normalize((0.485, 0.456, 0.406), (0.229, 0.224, 0.225)),
])
# a CUB-200-2011 test image, class index 50 (051.Horned_Grebe)
x = tf(Image.open('Horned_Grebe_0050_34561.jpg').convert('RGB')).unsqueeze(0)
with torch.no_grad():
order, family, species = model(x) # one logits tensor per level, coarsest first
print(species.argmax(-1).item(), species.softmax(-1).max().item()) # 50 0.8240
```
## Accuracy of the released checkpoints
Top-1 accuracy (%) stored in each file: the run's own finest-level accuracy on the dataset's test
split after the last epoch. The papers report the maximum over a learning-rate sweep, so their
tables differ. Per-file accuracy is in `manifest.csv`; the table below covers
`serial_21`.
| Model | aircraft (21) | cars (21) | cub (21) |
|---|---|---|---|
| `hideit3_base_patch16_224.fb_in1k` | 92.08 | 93.40 | 88.63 |
| `hideit3_base_patch16_224.fb_in22k_ft_in1k` | 92.98 | 93.65 | 90.63 |
| `hideit_base_patch16_224.fb_in1k` | 92.11 | 93.06 | 88.82 |
| `hiresnet50.a1_in1k` | 92.92 | 93.14 | 85.78 |
| `hiresnet50.fb_ssl_yfcc100m_ft_in1k` | 93.07 | 94.28 | 85.55 |
| `hiresnet50.fb_swsl_ig1b_ft_in1k` | 93.10 | 94.28 | 85.48 |
| `hiresnet50.gluon_in1k` | 93.01 | 94.23 | 85.24 |
| `hiresnet50.in1k_mocov3` | 93.07 | 93.53 | 84.85 |
| `hiresnet50.in1k_spark` | 91.27 | 92.38 | 78.63 |
| `hiresnet50.in1k_supcon` | 93.34 | 94.24 | 85.38 |
| `hiresnet50.in1k_swav` | 92.65 | 93.86 | 86.02 |
| `hiresnet50.in21k_miil` | 93.01 | 93.74 | 87.85 |
| `hiresnet50.tv2_in1k` | 92.80 | 93.81 | 87.28 |
| `hiresnet50.tv_in1k` | 92.71 | 94.09 | 86.16 |
| `hivit_base_patch16_224.dino` | 88.00 | 91.10 | 86.47 |
| `hivit_base_patch16_224.in1k_mocov3` | 90.19 | 93.07 | 87.73 |
| `hivit_base_patch16_224.mae` | 83.62 | 92.79 | 85.26 |
| `hivit_base_patch16_224.orig_in21k` | 89.32 | 91.73 | 90.97 |
| `hivit_base_patch16_224_miil.in21k` | 90.49 | 92.13 | 90.21 |
| `hivit_base_patch16_clip_224.laion2b` | 86.41 | 92.36 | 80.45 |
| `hivit_base_patch16_siglip_224.v2_webli` | 93.40 | 95.06 | 87.81 |
## Requirements
- `fgir-zoo` (`pip install git+https://github.com/arkel23/fgir-zoo.git`), which pins
`timm==0.9.12`
- `torch` (checked with 2.5.1)
## Citation
```bibtex
@inproceedings{surya_revisiting_2025,
title = {Revisiting the Backbone, Pretraining and Transferability for Hierarchical
Fine-Grained Image Recognition},
author = {Surya, Augusto Christian and Rios, Edwin Arkel and Lai, Bo-Cheng and Hu, Min-Chun},
booktitle = {Computer Vision, Graphics, and Image Processing (CVGIP)},
year = {2025},
eprint = {TBA},
note = {A. C. Surya and E. A. Rios contributed equally. arXiv ID to be added.}
}
```
|