Instructions to use emrekuruu/RetrievalRouter_Baseline with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use emrekuruu/RetrievalRouter_Baseline with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="emrekuruu/RetrievalRouter_Baseline", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("emrekuruu/RetrievalRouter_Baseline", trust_remote_code=True, device_map="auto") - PEFT
How to use emrekuruu/RetrievalRouter_Baseline with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
RetrievalRouter — Strategy-Selection Baseline
Adaptive routing baseline from RetrievalRouter: Joint Modality and Architecture Selection for Document Retrieval (EMNLP 2026). This is not the RetrievalRouter method — it is the prior strategy-selection approach of Arabzadeh et al. (CIKM 2021), extended to our five-pipeline setting and trained under an identical pipeline for a controlled comparison.
- 📄 Paper: https://arxiv.org/pdf/2608.25625
- 💻 Code: https://github.com/emrekuruu/retrieval-router
- 🤗 RetrievalRouter checkpoints: https://huggingface.co/collections/emrekuruu/retrieval-router
Motivation
Retrieval pipelines differ in modality (search over text, or over page images) and architecture (cheap dense search, or expensive late-interaction). The accurate ones are slow; the fast ones miss evidence on hard documents. And which one fails depends on the query — a text pipeline can't answer "what's the red curve in Figure 3?", but a multimodal one is overkill for a plain factoid. Across 11 benchmarks, no single pipeline wins on everything. RetrievalRouter picks the cheapest pipeline that can still answer each query, so easy queries stay fast and hard ones still get the heavy pipeline.
What this model is for
A research baseline, provided to reproduce the paper's comparisons. It trains a hard-label classifier: each query is labeled with the cheapest pipeline that ranks a relevant page first (BM25 when none do); at inference a probability threshold decides whether the query stays on BM25 or escalates to the top-scoring neural pipeline. It differs from RetrievalRouter only in the training signal — hard labels + cross-entropy here, versus a soft reward vector + KL for RetrievalRouter. Everything else (encoder, LoRA, pooling, head, split, optimizer, schedule) is shared, so any difference reflects the training signal alone. For deployment, prefer the RetrievalRouter checkpoints, which dominate this baseline on the accuracy–latency frontier.
Routing arms
| Index | Arm (config name) | Paper name | Modality | Architecture |
|---|---|---|---|---|
| 0 | MULTIMODAL_RERANK |
MM-Rerank | Multimodal | Dense → late-interaction rerank |
| 1 | MULTIMODAL-SINGLE |
MM-Dense | Multimodal | Single-vector dense |
| 2 | TEXT_RERANK |
Text-Rerank | Text | Dense → late-interaction rerank |
| 3 | TEXT-SINGLE |
Text-Dense | Text | Single-vector dense |
| 4 | BM25 |
BM25 | Text | Lexical |
Architecture
Identical to the RetrievalRouter checkpoints:
Qwen/Qwen3-0.6B-Base encoder + LoRA (merged), mean
pooling → 1024-d, single linear head over the five arms. Custom modeling code loads via
trust_remote_code=True.
Usage
import torch
from transformers import AutoModel, AutoTokenizer
repo = "emrekuruu/RetrievalRouter_Baseline"
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModel.from_pretrained(repo, trust_remote_code=True).eval()
inputs = tokenizer("In figure 3, what does the red dashed curve represent?",
return_tensors="pt", truncation=True, max_length=128)
with torch.no_grad():
logits = model(**inputs)["logits"] # shape [1, 5]
arm = model.config.strategy_names[logits.softmax(-1).argmax(-1).item()]
print(arm) # e.g. "MULTIMODAL_RERANK" -> run that pipeline for this query
The router returns which retrieval pipeline to run, not documents. You then execute the selected pipeline against your own indices.
Citation
If you use this baseline, please cite both the RetrievalRouter paper and the original strategy-selection method it is based on:
@misc{kuru2026retrievalrouterjointmodalityarchitecture,
title={RetrievalRouter: Joint Modality and Architecture Selection for Document Retrieval},
author={Emre Kuru and Mehmet Onur Keskin and Reza Farahbakhsh and Noel Crespi},
year={2026},
eprint={2608.25625},
archivePrefix={arXiv},
primaryClass={cs.IR},
url={https://arxiv.org/abs/2608.25625},
}
@inproceedings{arabzadeh2021predicting,
title = {Predicting efficiency/effectiveness trade-offs for dense vs. sparse retrieval strategy selection},
author = {Arabzadeh, Negar and Yan, Xinyi and Clarke, Charles L. A.},
booktitle = {Proceedings of the 30th ACM International Conference on Information \& Knowledge Management (CIKM)},
pages = {2862--2866},
year = {2021}
}
- Downloads last month
- 11
Model tree for emrekuruu/RetrievalRouter_Baseline
Base model
Qwen/Qwen3-0.6B-Base