bert-tiny-multi-vector / RUN_SUMMARY.md
tomaarsen's picture
tomaarsen HF Staff
Add new MultiVectorEncoder model
a0b72ca verified
|
Raw
History Blame
1.99 kB

Longer BERT tiny training run

Continued multi-vector-encoder-testing/bert-tiny-msmarco at revision 81c5b4e78ac3bdbb01606e60e82bc34d86ed897b.

Trained for 10,000 additional steps with batch size 128 on 501,907 MS MARCO BM25 triplets, with 1,024 held-out rows for evaluation loss. This processed 1.28 million triplets over about 2.55 epochs. Learning rate was 1e-5, with 5% warmup and linear decay. Training used bf16 on one RTX 3090 and took 15.6 minutes including periodic evaluation.

Selected step 6,000 using mean nDCG@10 on NanoMSMARCO, NanoNQ, and NanoFiQA2018, evaluated every 2,000 steps. The remaining ten datasets were evaluated only before and after training.

Mean nDCG@10

Evaluation group Initial model Longer run
Three selection datasets 0.3326 0.3729
All 13 datasets 0.3999 0.4468
Ten additional datasets 0.4201 0.4690

Per-dataset nDCG@10

Dataset Initial model Longer run Change
NanoClimateFEVER 0.1431 0.1919 +0.0489
NanoDBPedia 0.3914 0.4820 +0.0907
NanoFEVER 0.6310 0.6793 +0.0483
NanoFiQA2018 0.2658 0.3027 +0.0369
NanoHotpotQA 0.5766 0.6840 +0.1073
NanoMSMARCO 0.4328 0.3859 -0.0469
NanoNFCorpus 0.2569 0.2852 +0.0283
NanoNQ 0.2992 0.4300 +0.1307
NanoQuoraRetrieval 0.7976 0.8355 +0.0379
NanoSCIDOCS 0.2029 0.2217 +0.0188
NanoArguAna 0.2863 0.3539 +0.0676
NanoSciFact 0.5060 0.5650 +0.0590
NanoTouche2020 0.4096 0.3920 -0.0176

Improved on 11 of 13 datasets. NanoMSMARCO and NanoTouche2020 declined.

Reproduce

Run with a compatible Sentence Transformers checkout and its training dependencies:

python train.py --long-run

The script pins the initial model revision. Training arguments, every evaluation, and per-dataset before-and-after results are included in results.json.