Context reranker (MiniLM-L6, int8 ONNX)
A small cross-encoder that scores how well a passage answers a question. In a retrieval-augmented system it orders the parts of the retrieved documents the answer model reads, so the passage with the answer gets into a limited context budget first.
- Base model: cross-encoder/ms-marco-MiniLM-L-6-v2 (22M parameters)
- Teacher: BAAI/bge-reranker-v2-m3 — strong, but seconds per question on CPU
- Training: listwise distillation of the teacher's scores (KL divergence), 2 epochs, on about 9,000 English questions that a locally hosted LLM wrote from 3,000 passages of a public-finance document corpus. For each question, the teacher scored the passages of the top-10 sections the retrieval system returned for it. No evaluation questions were used for training.
- Format: ONNX (exported with classic attention, opset 14), dynamic int8 quantization, 23 MB. About 65 ms to score 16 passages on a desktop CPU.
Results
| Test | Base MiniLM | This model | Teacher |
|---|---|---|---|
| In-domain: answer passage inside an 18,000-character context (475 questions; 396 with no reranker; this model scored by copies that never saw the question's document) | 398 | 409 | 415 |
| SciFact, 100 questions, BM25 top-30 reranked (nDCG@10) | 0.723 | 0.757 | 0.793 |
| 20 everyday questions, answer ranked first of 4 | 17/20 | 17/20 | 18/20 |
Usage
With fastembed:
from huggingface_hub import snapshot_download
from fastembed.rerank.cross_encoder import TextCrossEncoder
from fastembed.common.model_description import ModelSource
path = snapshot_download("Badawi-a/context-reranker", revision="v1")
TextCrossEncoder.add_custom_model(model="local/context-reranker", sources=ModelSource(hf="local/context-reranker"),
model_file="onnx/model.onnx")
reranker = TextCrossEncoder("local/context-reranker", specific_model_path=path)
scores = list(reranker.rerank("What is the VAT rate?", ["The VAT rate is 14 percent.", "It rained in Cairo."], batch_size=8))
Higher scores mean more relevant.
Limitations
- English only. Questions in other languages should be translated first.
- Trained for one domain (public finance); it generalizes (see SciFact), but less well than the teacher.
- Like other small rerankers, it can prefer a passage that repeats the question's words over one that answers it.
- Downloads last month
- 17
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for Badawi-a/context-reranker
Base model
microsoft/MiniLM-L12-H384-uncased Quantized
cross-encoder/ms-marco-MiniLM-L12-v2 Quantized
cross-encoder/ms-marco-MiniLM-L6-v2