# Longer BERT tiny training run Continued multi-vector-encoder-testing/bert-tiny-msmarco at revision 81c5b4e78ac3bdbb01606e60e82bc34d86ed897b. Trained for 10,000 additional steps with batch size 128 on 501,907 MS MARCO BM25 triplets, with 1,024 held-out rows for evaluation loss. This processed 1.28 million triplets over about 2.55 epochs. Learning rate was 1e-5, with 5% warmup and linear decay. Training used bf16 on one RTX 3090 and took 15.6 minutes including periodic evaluation. Selected step 6,000 using mean nDCG@10 on NanoMSMARCO, NanoNQ, and NanoFiQA2018, evaluated every 2,000 steps. The remaining ten datasets were evaluated only before and after training. ## Mean nDCG@10 | Evaluation group | Initial model | Longer run | | --- | ---: | ---: | | Three selection datasets | 0.3326 | 0.3729 | | All 13 datasets | 0.3999 | 0.4468 | | Ten additional datasets | 0.4201 | 0.4690 | ## Per-dataset nDCG@10 | Dataset | Initial model | Longer run | Change | | --- | ---: | ---: | ---: | | NanoClimateFEVER | 0.1431 | 0.1919 | +0.0489 | | NanoDBPedia | 0.3914 | 0.4820 | +0.0907 | | NanoFEVER | 0.6310 | 0.6793 | +0.0483 | | NanoFiQA2018 | 0.2658 | 0.3027 | +0.0369 | | NanoHotpotQA | 0.5766 | 0.6840 | +0.1073 | | NanoMSMARCO | 0.4328 | 0.3859 | -0.0469 | | NanoNFCorpus | 0.2569 | 0.2852 | +0.0283 | | NanoNQ | 0.2992 | 0.4300 | +0.1307 | | NanoQuoraRetrieval | 0.7976 | 0.8355 | +0.0379 | | NanoSCIDOCS | 0.2029 | 0.2217 | +0.0188 | | NanoArguAna | 0.2863 | 0.3539 | +0.0676 | | NanoSciFact | 0.5060 | 0.5650 | +0.0590 | | NanoTouche2020 | 0.4096 | 0.3920 | -0.0176 | Improved on 11 of 13 datasets. NanoMSMARCO and NanoTouche2020 declined. ## Reproduce Run with a compatible Sentence Transformers checkout and its training dependencies: ```bash python train.py --long-run ``` The script pins the initial model revision. Training arguments, every evaluation, and per-dataset before-and-after results are included in results.json.