RetrievalRouter — Strategy-Selection Baseline

Adaptive routing baseline from RetrievalRouter: Joint Modality and Architecture Selection for Document Retrieval (EMNLP 2026). This is not the RetrievalRouter method — it is the prior strategy-selection approach of Arabzadeh et al. (CIKM 2021), extended to our five-pipeline setting and trained under an identical pipeline for a controlled comparison.

Motivation

Retrieval pipelines differ in modality (search over text, or over page images) and architecture (cheap dense search, or expensive late-interaction). The accurate ones are slow; the fast ones miss evidence on hard documents. And which one fails depends on the query — a text pipeline can't answer "what's the red curve in Figure 3?", but a multimodal one is overkill for a plain factoid. Across 11 benchmarks, no single pipeline wins on everything. RetrievalRouter picks the cheapest pipeline that can still answer each query, so easy queries stay fast and hard ones still get the heavy pipeline.

What this model is for

A research baseline, provided to reproduce the paper's comparisons. It trains a hard-label classifier: each query is labeled with the cheapest pipeline that ranks a relevant page first (BM25 when none do); at inference a probability threshold decides whether the query stays on BM25 or escalates to the top-scoring neural pipeline. It differs from RetrievalRouter only in the training signal — hard labels + cross-entropy here, versus a soft reward vector + KL for RetrievalRouter. Everything else (encoder, LoRA, pooling, head, split, optimizer, schedule) is shared, so any difference reflects the training signal alone. For deployment, prefer the RetrievalRouter checkpoints, which dominate this baseline on the accuracy–latency frontier.

Routing arms

Index Arm (config name) Paper name Modality Architecture
0 MULTIMODAL_RERANK MM-Rerank Multimodal Dense → late-interaction rerank
1 MULTIMODAL-SINGLE MM-Dense Multimodal Single-vector dense
2 TEXT_RERANK Text-Rerank Text Dense → late-interaction rerank
3 TEXT-SINGLE Text-Dense Text Single-vector dense
4 BM25 BM25 Text Lexical

Architecture

Identical to the RetrievalRouter checkpoints: Qwen/Qwen3-0.6B-Base encoder + LoRA (merged), mean pooling → 1024-d, single linear head over the five arms. Custom modeling code loads via trust_remote_code=True.

Usage

import torch
from transformers import AutoModel, AutoTokenizer

repo = "emrekuruu/RetrievalRouter_Baseline"
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModel.from_pretrained(repo, trust_remote_code=True).eval()

inputs = tokenizer("In figure 3, what does the red dashed curve represent?",
                   return_tensors="pt", truncation=True, max_length=128)
with torch.no_grad():
    logits = model(**inputs)["logits"]              # shape [1, 5]
arm = model.config.strategy_names[logits.softmax(-1).argmax(-1).item()]
print(arm)   # e.g. "MULTIMODAL_RERANK" -> run that pipeline for this query

The router returns which retrieval pipeline to run, not documents. You then execute the selected pipeline against your own indices.

Citation

If you use this baseline, please cite both the RetrievalRouter paper and the original strategy-selection method it is based on:

@misc{kuru2026retrievalrouterjointmodalityarchitecture,
      title={RetrievalRouter: Joint Modality and Architecture Selection for Document Retrieval}, 
      author={Emre Kuru and Mehmet Onur Keskin and Reza Farahbakhsh and Noel Crespi},
      year={2026},
      eprint={2608.25625},
      archivePrefix={arXiv},
      primaryClass={cs.IR},
      url={https://arxiv.org/abs/2608.25625}, 
}

@inproceedings{arabzadeh2021predicting,
  title     = {Predicting efficiency/effectiveness trade-offs for dense vs. sparse retrieval strategy selection},
  author    = {Arabzadeh, Negar and Yan, Xinyi and Clarke, Charles L. A.},
  booktitle = {Proceedings of the 30th ACM International Conference on Information \& Knowledge Management (CIKM)},
  pages     = {2862--2866},
  year      = {2021}
}
Downloads last month
11
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for emrekuruu/RetrievalRouter_Baseline

Adapter
(68)
this model

Datasets used to train emrekuruu/RetrievalRouter_Baseline

Collection including emrekuruu/RetrievalRouter_Baseline

Paper for emrekuruu/RetrievalRouter_Baseline