Instructions to use multi-vector-encoder-testing/bert-tiny-multi-vector with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use multi-vector-encoder-testing/bert-tiny-multi-vector with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("multi-vector-encoder-testing/bert-tiny-multi-vector") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
Longer BERT tiny training run
Continued multi-vector-encoder-testing/bert-tiny-msmarco at revision 81c5b4e78ac3bdbb01606e60e82bc34d86ed897b.
Trained for 10,000 additional steps with batch size 128 on 501,907 MS MARCO BM25 triplets, with 1,024 held-out rows for evaluation loss. This processed 1.28 million triplets over about 2.55 epochs. Learning rate was 1e-5, with 5% warmup and linear decay. Training used bf16 on one RTX 3090 and took 15.6 minutes including periodic evaluation.
Selected step 6,000 using mean nDCG@10 on NanoMSMARCO, NanoNQ, and NanoFiQA2018, evaluated every 2,000 steps. The remaining ten datasets were evaluated only before and after training.
Mean nDCG@10
| Evaluation group | Initial model | Longer run |
|---|---|---|
| Three selection datasets | 0.3326 | 0.3729 |
| All 13 datasets | 0.3999 | 0.4468 |
| Ten additional datasets | 0.4201 | 0.4690 |
Per-dataset nDCG@10
| Dataset | Initial model | Longer run | Change |
|---|---|---|---|
| NanoClimateFEVER | 0.1431 | 0.1919 | +0.0489 |
| NanoDBPedia | 0.3914 | 0.4820 | +0.0907 |
| NanoFEVER | 0.6310 | 0.6793 | +0.0483 |
| NanoFiQA2018 | 0.2658 | 0.3027 | +0.0369 |
| NanoHotpotQA | 0.5766 | 0.6840 | +0.1073 |
| NanoMSMARCO | 0.4328 | 0.3859 | -0.0469 |
| NanoNFCorpus | 0.2569 | 0.2852 | +0.0283 |
| NanoNQ | 0.2992 | 0.4300 | +0.1307 |
| NanoQuoraRetrieval | 0.7976 | 0.8355 | +0.0379 |
| NanoSCIDOCS | 0.2029 | 0.2217 | +0.0188 |
| NanoArguAna | 0.2863 | 0.3539 | +0.0676 |
| NanoSciFact | 0.5060 | 0.5650 | +0.0590 |
| NanoTouche2020 | 0.4096 | 0.3920 | -0.0176 |
Improved on 11 of 13 datasets. NanoMSMARCO and NanoTouche2020 declined.
Reproduce
Run with a compatible Sentence Transformers checkout and its training dependencies:
python train.py --long-run
The script pins the initial model revision. Training arguments, every evaluation, and per-dataset before-and-after results are included in results.json.