Visual Document Retrieval
Transformers
Safetensors
multilingual
qwen3_5
feature-extraction
text
image
multimodal-embedding
vidore
colbert
colqwen3_5
multilingual-embedding
custom_code
Instructions to use webAI-Official/webAI-ColVec1.1-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use webAI-Official/webAI-ColVec1.1-4b with Transformers:
# Load model directly from transformers import AutoProcessor, AutoModel processor = AutoProcessor.from_pretrained("webAI-Official/webAI-ColVec1.1-4b", trust_remote_code=True) model = AutoModel.from_pretrained("webAI-Official/webAI-ColVec1.1-4b", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 10,578 Bytes
e8491df b4f97a7 e8491df b4f97a7 e8491df 4332508 b4f97a7 e8491df 8873615 e8491df 8873615 e8491df b4f97a7 e8491df b4f97a7 e8491df | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 | ---
pipeline_tag: visual-document-retrieval
library_name: transformers
language:
- multilingual
license: other
license_name: webai-non-commercial-license-v1.0
license_link: https://huggingface.co/webAI-Official/webAI-ColVec1.1-4b/blob/main/LICENSE.md
base_model: Qwen/Qwen3.5-4B
datasets:
- vidore/colpali_train_set
- Tevatron/docmatix-ir
- openbmb/VisRAG-Ret-Train-In-domain-data
- openbmb/VisRAG-Ret-Train-Synthetic-data
- llamaindex/vdr-multilingual-train
- Tevatron/wiki-ss-nq
tags:
- text
- image
- multimodal-embedding
- visual-document-retrieval
- vidore
- colbert
- colqwen3_5
- multilingual-embedding
---
# webAI-Official/webAI-ColVec1.1-4b
## ⚡ Summary
**webAI-Official/webAI-ColVec1.1-4b** is a ColBERT-style multimodal
embedding model based on
[Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B). It maps text
queries and visual documents (images or rendered PDF pages) into aligned,
L2-normalized multi-vector embeddings for late-interaction retrieval.
The model uses bidirectional attention in Qwen3.5's full-attention layers and
a learned 640-dimensional projection head. The unused language-model head has
been removed from the released checkpoint.
### Training data
We created filtered, balanced, and multilingual curated subsets from six public
datasets: [ColPali Train Set](https://huggingface.co/datasets/vidore/colpali_train_set),
[Docmatix-IR](https://huggingface.co/datasets/Tevatron/docmatix-ir),
[VisRAG In-Domain](https://huggingface.co/datasets/openbmb/VisRAG-Ret-Train-In-domain-data),
[VisRAG Synthetic](https://huggingface.co/datasets/openbmb/VisRAG-Ret-Train-Synthetic-data),
[VDR Multilingual Train](https://huggingface.co/datasets/llamaindex/vdr-multilingual-train),
and [Wiki-SS-NQ](https://huggingface.co/datasets/Tevatron/wiki-ss-nq). This
Qwen3.5-4B-backbone model was trained on a 500,000-sample curated subset as well
as synthetically generated data.
## 🛠️ Model specifications
| Feature | Detail |
| :--- | :--- |
| **Architecture** | Qwen3.5-4B vision-language model + 640-dimensional linear projection |
| **Released parameters** | 4,540,904,576 |
| **Method** | ColBERT-style late interaction with MaxSim scoring |
| **Output** | L2-normalized multi-vector embeddings `(sequence_length, 640)` |
| **Modalities** | Text queries and document images |
| **Attention** | Bidirectional full-attention layers; selectable FlashAttention 2, FlashAttention 3, or SDPA kernel |
| **Visual-token budget** | 1,792 tokens per image in the released processor |
| **Training** | LoRA adapters and a fully trained projection layer, merged for release |
| **Weights** | `bfloat16`; language-model head removed |
### Key properties
- **Unified encoder:** The same model encodes text and document images.
- **Token-level retrieval:** Multi-vector embeddings preserve fine-grained
layout and content signals that single-vector pooling can discard.
- **Compact projection:** Hidden states are projected to 640 dimensions
without an activation function.
- **Bidirectional retrieval attention:** Selecting FlashAttention 2 or 3
changes the execution kernel, not the model's bidirectional attention mode.
## 📊 Evaluation results
The table reports NDCG@10 scores on the eight **public** ViDoRe V3 tasks as
percentages rather than values between 0 and 1 (for example, 0.80 is shown as
80.00). Each task value is the mean of its six language subsets; the public
average is the unweighted mean of the eight public task values.
The result artifacts record MTEB 2.18.5, Transformers 5.14.1, PyTorch 2.9.0
with CUDA 12.8, `bfloat16`, FlashAttention 2.8.3, and batch size 32. The release
processor is configured with a 1,792 visual-token budget. Comparator values
were read from the live
[ViDoRe V3 MTEB leaderboard](https://mteb-leaderboard.hf.space/benchmark/ViDoRe%28v3%29)
on July 22, 2026.
Model encoding runs in `bfloat16`. Before MaxSim scoring, query and document
embeddings are moved to CPU and converted to `float32`. All reported ViDoRe
results use this FP32 scoring path.
| Model | Computer Science | Energy | FinanceEn | FinanceFr | HR | Industrial | Pharmaceuticals | Physics | **Avg. public** |
| :--- | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: |
| **[webAI-ColVec1.1-8b](https://huggingface.co/webAI-Official/webAI-ColVec1.1-8b)** | 80.14 | 70.09 | **71.89** | **55.10** | 68.49 | **57.41** | 67.88 | 51.29 | **65.29** |
| [VultronRetriever Prime](https://huggingface.co/vultr/VultronRetrieverPrime-Qwen3.5-8B) | 79.81 | **70.26** | 69.01 | 54.51 | 66.82 | **57.41** | **68.19** | **51.73** | 64.72 |
| [webAI-ColVec1-9b](https://huggingface.co/webAI-Official/webAI-ColVec1-9b) | **80.92** | 69.77 | 68.28 | 53.72 | **70.04** | 57.18 | 67.32 | 48.38 | 64.45 |
| **webAI-ColVec1.1-4b (this model)** | 80.35 | 69.26 | 69.12 | 53.21 | 67.02 | 56.30 | 67.07 | 51.36 | 64.21 |
| [VultronRetriever Core](https://huggingface.co/vultr/VultronRetrieverCore-Qwen3.5-4.5B) | 79.77 | 69.19 | 68.93 | 52.02 | 66.10 | 56.11 | 67.45 | 50.18 | 63.72 |
| [Nemotron ColEmbed VL 8B V2](https://huggingface.co/nvidia/nemotron-colembed-vl-8b-v2) | 79.29 | 69.82 | 67.29 | 51.54 | 66.32 | 56.03 | 67.19 | 50.84 | 63.54 |
| [webAI-ColVec1-4b](https://huggingface.co/webAI-Official/webAI-ColVec1-4b) | 79.84 | 68.70 | 68.49 | 51.11 | 67.40 | 55.73 | 65.68 | 50.15 | 63.39 |
| [Tomoro ColQwen3 Embed 8B](https://huggingface.co/TomoroAI/tomoro-colqwen3-embed-8b) | 75.35 | 68.41 | 65.08 | 49.10 | 63.98 | 54.41 | 66.36 | 50.13 | 61.60 |
The current MTEB leaderboard entries named `webAI-ColVec1-4b` and
`webAI-ColVec1-9b` refer to the previous ColVec1 release, not these ColVec1.1
checkpoints.
## 💻 Usage
The processor provides the current retrieval API:
- `process_images(images)` prepares one or more document images.
- `process_queries(texts)` prepares one or more natural-language queries.
- `score_retrieval(query_embeddings, document_embeddings)` computes a MaxSim
score matrix with shape `(number_of_queries, number_of_documents)`.
### Prerequisites and attention backends
The public ViDoRe V3 numbers are reproducible with this pinned FlashAttention 2
environment:
```text
Python 3.12
PyTorch 2.9.0 + CUDA 12.8
Transformers 5.14.1
MTEB 2.18.5
Sentence Transformers 5.6.0
FlashAttention 2.8.3
```
The model also supports FlashAttention 3 on Hopper GPUs (H100/H200). See the
[Dao-AILab FlashAttention repository](https://github.com/dao-ailab/flash-attention)
for FlashAttention 3 installation instructions, then select
`flash_attention_3` when loading in a compatible Hopper environment.
FlashAttention 2 and FlashAttention 3 both preserve bidirectional attention.
FlashAttention 3 can improve throughput on Hopper GPUs, but the published
scores use FlashAttention 2; changing kernels can produce small floating-point
differences. Use FlashAttention 2 when reproducing the table.
### Inference code
```python
from io import BytesIO
import requests
import torch
from PIL import Image
from transformers import AutoModel, AutoProcessor
MODEL_ID = "webAI-Official/webAI-ColVec1.1-4b"
DEVICE = "cuda:0" if torch.cuda.is_available() else "cpu"
ATTN_IMPLEMENTATION = (
"flash_attention_2" if torch.cuda.is_available() else "sdpa"
)
# On an H100/H200 with FlashAttention 3 installed, use:
# ATTN_IMPLEMENTATION = "flash_attention_3"
processor = AutoProcessor.from_pretrained(
MODEL_ID,
trust_remote_code=True,
max_num_visual_tokens=1792,
)
model = AutoModel.from_pretrained(
MODEL_ID,
trust_remote_code=True,
dtype=torch.bfloat16,
attn_implementation=ATTN_IMPLEMENTATION,
device_map=DEVICE,
).eval()
queries = [
"When was the United States Declaration of Independence proclaimed?",
"Who printed the edition of Romeo and Juliet?",
]
document_urls = [
"https://upload.wikimedia.org/wikipedia/commons/8/89/US-original-Declaration-1776.jpg",
"https://upload.wikimedia.org/wikipedia/commons/thumb/4/4c/Romeoandjuliet1597.jpg/500px-Romeoandjuliet1597.jpg",
]
def load_image(url: str) -> Image.Image:
response = requests.get(
url,
headers={"User-Agent": "Mozilla/5.0"},
timeout=30,
)
response.raise_for_status()
return Image.open(BytesIO(response.content)).convert("RGB")
device = next(model.parameters()).device
query_inputs = processor.process_queries(queries)
document_inputs = processor.process_images(
[load_image(url) for url in document_urls]
)
query_inputs = {
key: value.to(device) if isinstance(value, torch.Tensor) else value
for key, value in query_inputs.items()
}
document_inputs = {
key: value.to(device) if isinstance(value, torch.Tensor) else value
for key, value in document_inputs.items()
}
with torch.inference_mode():
query_batch = model(**query_inputs)
document_batch = model(**document_inputs)
query_embeddings = [embedding.cpu() for embedding in query_batch]
document_embeddings = [embedding.cpu() for embedding in document_batch]
scores = processor.score_retrieval(
query_embeddings,
document_embeddings,
output_dtype=torch.float32,
)
print(scores)
print("Best document per query:", scores.argmax(dim=1))
```
The processor loads the release's 1,792 visual-token budget by default. To
reduce memory use, pass a lower `max_num_visual_tokens` value to
`AutoProcessor.from_pretrained`; this changes document granularity and may
change retrieval scores.
## ⚖️ Strengths and limitations
### Strengths
- **Performance:** State-of-the-art retrieval performance among 4B models on
the public ViDoRe V3 tasks, with excellent multimodal document retrieval
results.
- **Complex layouts:** Excellent handling of chart-rich PDFs and
domain-specific documents.
- **End-to-end retrieval:** OCR-free retrieval on unseen multimodal documents
without using an intermediate vision-language model to generate summaries.
- **Multilingualism:** Strong performance on non-English document inputs.
### Limitations
- **Storage Cost:** Still larger than single-vector baselines despite the
smaller token dimension.
## License
Model weights are distributed under the
[webAI Non-Commercial License v1.0](https://huggingface.co/webAI-Official/webAI-ColVec1.1-4b/blob/main/LICENSE.md).
See the repository's `NOTICES.md` for upstream attribution.
## 📚 Citation
```bibtex
@misc{webai_colvec1_1_4b,
title = {webAI-ColVec1.1-4b: A Bidirectional Multi-Vector Model for Visual Document Retrieval},
author = {webAI},
year = {2026},
url = {https://huggingface.co/webAI-Official/webAI-ColVec1.1-4b}
}
```
|