Instructions to use ERISLab/Hierarchical-Backbones with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- timm
How to use ERISLab/Hierarchical-Backbones with timm:
import timm model = timm.create_model("hf_hub:ERISLab/Hierarchical-Backbones", pretrained=True) - Notebooks
- Google Colab
- Kaggle
|
Download README.md from ERISLab/Hierarchical-Backbones: direct link, hf CLI and curl.
- Browser
- Download file 6 kB
-
https://huggingface.co/ERISLab/Hierarchical-Backbones/resolve/main/README.md
- Command line
-
hf download hf://ERISLab/Hierarchical-Backbones/README.md
-
curl -L -o README.md https://huggingface.co/ERISLab/Hierarchical-Backbones/resolve/main/README.md
6 kB
| pipeline_tag: image-classification | |
| library_name: pytorch | |
| tags: | |
| - fine-grained-image-recognition | |
| - image-classification | |
| - hierarchical-classification | |
| - timm | |
| # Hierarchical-Backbones: hierarchical classification across pretrained backbones | |
| These are the checkpoints behind *Revisiting the Backbone, Pretraining and Transferability for Hierarchical Fine-Grained Image Recognition*, CVGIP 2025 (arXiv link to be added). Twenty-one pretrained backbones (ResNet-50, ViT-B, DeiT-B and DeiT-3-B under supervised, self-supervised and vision-language pretraining) fine-tuned for fine-grained recognition with a hierarchical classification head. The head predicts every level of the annotated label hierarchy, coarsest first: three levels on CUB and Aircraft, two on Cars. Code: [arkel23/Hierarchical](https://github.com/arkel23/Hierarchical). | |
| 168 checkpoints, one per configuration, each the last epoch of one training run at 448 px. | |
| Each file is a `torch.save` dict with `config` (the full training configuration), `model` (the | |
| state dict of the backbone and head), `accuracy` and `epoch`, with no optimizer state. File names | |
| are the runs' experiment-log names: dataset, cluster ratio for a pseudo-hierarchy, model, `fz` for | |
| a frozen backbone, and the serial. Load them with | |
| [fgir-zoo](https://github.com/arkel23/fgir-zoo). The [collection](https://huggingface.co/collections/ERISLab/hierarchical-fgir-backbones-and-pretraining-cvgip-2025-6abab163857f4e028b3a38a1) is this work's | |
| page on the Hub; the paper joins it once it is on arXiv. | |
| ## Layout | |
| One folder per serial. | |
| | Folder | What | Files | Mean accuracy | | |
| |---|---|---|---| | |
| | `serial_21` | Real (annotated) hierarchy, full fine-tuning, CUB/Aircraft/Cars x 21 backbones, 200 epochs, each backbone at its best learning rate from the paper's sweep, seed 1 | 63 | 90.39 | | |
| | `serial_58` | Aircraft x 21 backbones, 50 epochs, best learning rate, with the real hierarchy; matched with `serial_59` | 21 | 91.39 | | |
| | `serial_59` | Aircraft x 21 backbones, 50 epochs, best learning rate, without a hierarchy; matched with `serial_58` | 21 | 90.94 | | |
| | `serial_60` | Real hierarchy with a frozen backbone (only the classification head trained), the same 63 configurations at their best learning rate, 50 or 200 epochs | 63 | 64.90 | | |
| `manifest.csv` lists every file with its dataset, model, hierarchy (`real`, `pseudo` or `none`), | |
| cluster ratio, frozen backbone, serial, seed, epochs, image size, class count, number of levels, | |
| accuracy, SHA-256 and size. | |
| ## Load a checkpoint and classify an image | |
| With a hierarchy the model returns a tuple of logits, one per level from coarsest to finest; | |
| without one it returns a single logits tensor. The finest level is the prediction. The evaluation | |
| transform is the training code's: a bicubic resize to 550 x 550, a 448 px center crop and ImageNet | |
| normalization. | |
| ```python | |
| import torch | |
| from PIL import Image | |
| from torchvision import transforms | |
| from fgir_zoo import hierarchical | |
| model = hierarchical.create_model('serial_21/cub_hideit_base_patch16_224.fb_in1k_21') | |
| cfg = model.config | |
| tf = transforms.Compose([ | |
| transforms.Resize((cfg.test_resize_size, cfg.test_resize_size), | |
| interpolation=transforms.InterpolationMode.BICUBIC), | |
| transforms.CenterCrop(cfg.image_size), | |
| transforms.ToTensor(), | |
| transforms.Normalize((0.485, 0.456, 0.406), (0.229, 0.224, 0.225)), | |
| ]) | |
| # a CUB-200-2011 test image, class index 50 (051.Horned_Grebe) | |
| x = tf(Image.open('Horned_Grebe_0050_34561.jpg').convert('RGB')).unsqueeze(0) | |
| with torch.no_grad(): | |
| order, family, species = model(x) # one logits tensor per level, coarsest first | |
| print(species.argmax(-1).item(), species.softmax(-1).max().item()) # 50 0.8240 | |
| ``` | |
| ## Accuracy of the released checkpoints | |
| Top-1 accuracy (%) stored in each file: the run's own finest-level accuracy on the dataset's test | |
| split after the last epoch. The papers report the maximum over a learning-rate sweep, so their | |
| tables differ. Per-file accuracy is in `manifest.csv`; the table below covers | |
| `serial_21`. | |
| | Model | aircraft (21) | cars (21) | cub (21) | | |
| |---|---|---|---| | |
| | `hideit3_base_patch16_224.fb_in1k` | 92.08 | 93.40 | 88.63 | | |
| | `hideit3_base_patch16_224.fb_in22k_ft_in1k` | 92.98 | 93.65 | 90.63 | | |
| | `hideit_base_patch16_224.fb_in1k` | 92.11 | 93.06 | 88.82 | | |
| | `hiresnet50.a1_in1k` | 92.92 | 93.14 | 85.78 | | |
| | `hiresnet50.fb_ssl_yfcc100m_ft_in1k` | 93.07 | 94.28 | 85.55 | | |
| | `hiresnet50.fb_swsl_ig1b_ft_in1k` | 93.10 | 94.28 | 85.48 | | |
| | `hiresnet50.gluon_in1k` | 93.01 | 94.23 | 85.24 | | |
| | `hiresnet50.in1k_mocov3` | 93.07 | 93.53 | 84.85 | | |
| | `hiresnet50.in1k_spark` | 91.27 | 92.38 | 78.63 | | |
| | `hiresnet50.in1k_supcon` | 93.34 | 94.24 | 85.38 | | |
| | `hiresnet50.in1k_swav` | 92.65 | 93.86 | 86.02 | | |
| | `hiresnet50.in21k_miil` | 93.01 | 93.74 | 87.85 | | |
| | `hiresnet50.tv2_in1k` | 92.80 | 93.81 | 87.28 | | |
| | `hiresnet50.tv_in1k` | 92.71 | 94.09 | 86.16 | | |
| | `hivit_base_patch16_224.dino` | 88.00 | 91.10 | 86.47 | | |
| | `hivit_base_patch16_224.in1k_mocov3` | 90.19 | 93.07 | 87.73 | | |
| | `hivit_base_patch16_224.mae` | 83.62 | 92.79 | 85.26 | | |
| | `hivit_base_patch16_224.orig_in21k` | 89.32 | 91.73 | 90.97 | | |
| | `hivit_base_patch16_224_miil.in21k` | 90.49 | 92.13 | 90.21 | | |
| | `hivit_base_patch16_clip_224.laion2b` | 86.41 | 92.36 | 80.45 | | |
| | `hivit_base_patch16_siglip_224.v2_webli` | 93.40 | 95.06 | 87.81 | | |
| ## Requirements | |
| - `fgir-zoo` (`pip install git+https://github.com/arkel23/fgir-zoo.git`), which pins | |
| `timm==0.9.12` | |
| - `torch` (checked with 2.5.1) | |
| ## Citation | |
| ```bibtex | |
| @inproceedings{surya_revisiting_2025, | |
| title = {Revisiting the Backbone, Pretraining and Transferability for Hierarchical | |
| Fine-Grained Image Recognition}, | |
| author = {Surya, Augusto Christian and Rios, Edwin Arkel and Lai, Bo-Cheng and Hu, Min-Chun}, | |
| booktitle = {Computer Vision, Graphics, and Image Processing (CVGIP)}, | |
| year = {2025}, | |
| eprint = {TBA}, | |
| note = {A. C. Surya and E. A. Rios contributed equally. arXiv ID to be added.} | |
| } | |
| ``` | |