Context reranker (MiniLM-L6, int8 ONNX)

A small cross-encoder that scores how well a passage answers a question. In a retrieval-augmented system it orders the parts of the retrieved documents the answer model reads, so the passage with the answer gets into a limited context budget first.

  • Base model: cross-encoder/ms-marco-MiniLM-L-6-v2 (22M parameters)
  • Teacher: BAAI/bge-reranker-v2-m3 — strong, but seconds per question on CPU
  • Training: listwise distillation of the teacher's scores (KL divergence), 2 epochs, on about 9,000 English questions that a locally hosted LLM wrote from 3,000 passages of a public-finance document corpus. For each question, the teacher scored the passages of the top-10 sections the retrieval system returned for it. No evaluation questions were used for training.
  • Format: ONNX (exported with classic attention, opset 14), dynamic int8 quantization, 23 MB. About 65 ms to score 16 passages on a desktop CPU.

Results

Test Base MiniLM This model Teacher
In-domain: answer passage inside an 18,000-character context (475 questions; 396 with no reranker; this model scored by copies that never saw the question's document) 398 409 415
SciFact, 100 questions, BM25 top-30 reranked (nDCG@10) 0.723 0.757 0.793
20 everyday questions, answer ranked first of 4 17/20 17/20 18/20

Usage

With fastembed:

from huggingface_hub import snapshot_download
from fastembed.rerank.cross_encoder import TextCrossEncoder
from fastembed.common.model_description import ModelSource

path = snapshot_download("Badawi-a/context-reranker", revision="v1")
TextCrossEncoder.add_custom_model(model="local/context-reranker", sources=ModelSource(hf="local/context-reranker"),
                                  model_file="onnx/model.onnx")
reranker = TextCrossEncoder("local/context-reranker", specific_model_path=path)
scores = list(reranker.rerank("What is the VAT rate?", ["The VAT rate is 14 percent.", "It rained in Cairo."], batch_size=8))

Higher scores mean more relevant.

Limitations

  • English only. Questions in other languages should be translated first.
  • Trained for one domain (public finance); it generalizes (see SciFact), but less well than the teacher.
  • Like other small rerankers, it can prefer a passage that repeats the question's words over one that answers it.
Downloads last month
17
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Badawi-a/context-reranker