webAI-ColVec1.1-8b / README.md
Zhan-Lun's picture
Update 8B model card
5a6eef0 verified
|
Raw
History Blame Contribute Delete
10.7 kB
---
pipeline_tag: visual-document-retrieval
library_name: transformers
language:
- multilingual
license: other
license_name: webai-non-commercial-license-v1.0
license_link: https://huggingface.co/webAI-Official/webAI-ColVec1.1-8b/blob/main/LICENSE.md
base_model: Qwen/Qwen3.5-9B
datasets:
- vidore/colpali_train_set
- Tevatron/docmatix-ir
- openbmb/VisRAG-Ret-Train-In-domain-data
- openbmb/VisRAG-Ret-Train-Synthetic-data
- llamaindex/vdr-multilingual-train
- Tevatron/wiki-ss-nq
tags:
- text
- image
- multimodal-embedding
- visual-document-retrieval
- vidore
- colbert
- colqwen3_5
- multilingual-embedding
---
# webAI-Official/webAI-ColVec1.1-8b
## ⚡ Summary
**webAI-Official/webAI-ColVec1.1-8b** is a ColBERT-style multimodal
embedding model based on
[Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B). It maps text
queries and visual documents (images or rendered PDF pages) into aligned,
L2-normalized multi-vector embeddings for late-interaction retrieval.
The model uses bidirectional attention in Qwen3.5's full-attention layers and
a learned 640-dimensional projection head. The unused language-model head has
been removed from the released checkpoint. Consequently, the released
embedding model contains ~8.4B parameters, reflected by the
`webAI-ColVec1.1-8b` repository name, while its original backbone is
Qwen3.5-9B.
### Training data
We created filtered, balanced, and multilingual curated subsets from six public
datasets: [ColPali Train Set](https://huggingface.co/datasets/vidore/colpali_train_set),
[Docmatix-IR](https://huggingface.co/datasets/Tevatron/docmatix-ir),
[VisRAG In-Domain](https://huggingface.co/datasets/openbmb/VisRAG-Ret-Train-In-domain-data),
[VisRAG Synthetic](https://huggingface.co/datasets/openbmb/VisRAG-Ret-Train-Synthetic-data),
[VDR Multilingual Train](https://huggingface.co/datasets/llamaindex/vdr-multilingual-train),
and [Wiki-SS-NQ](https://huggingface.co/datasets/Tevatron/wiki-ss-nq). This
Qwen3.5-9B-backbone model was trained on a 750,000-sample curated subset as well
as synthetically generated data.
## 🛠️ Model specifications
| Feature | Detail |
| :--- | :--- |
| **Architecture** | Qwen3.5-9B vision-language model + 640-dimensional linear projection |
| **Released parameters** | 8,395,317,104 |
| **Method** | ColBERT-style late interaction with MaxSim scoring |
| **Output** | L2-normalized multi-vector embeddings `(sequence_length, 640)` |
| **Modalities** | Text queries and document images |
| **Attention** | Bidirectional full-attention layers; selectable FlashAttention 2, FlashAttention 3, or SDPA kernel |
| **Visual-token budget** | 1,792 tokens per image in the released processor |
| **Training** | LoRA adapters and a fully trained projection layer, merged for release |
| **Weights** | `bfloat16`; language-model head removed |
### Key properties
- **Unified encoder:** The same model encodes text and document images.
- **Token-level retrieval:** Multi-vector embeddings preserve fine-grained
layout and content signals that single-vector pooling can discard.
- **Compact projection:** Hidden states are projected to 640 dimensions
without an activation function.
- **Bidirectional retrieval attention:** Selecting FlashAttention 2 or 3
changes the execution kernel, not the model's bidirectional attention mode.
## 📊 Evaluation results
The table reports NDCG@10 scores on the eight **public** ViDoRe V3 tasks as
percentages rather than values between 0 and 1 (for example, 0.80 is shown as
80.00). Each task value is the mean of its six language subsets; the public average is the unweighted mean of the eight public task values.
The result artifacts record MTEB 2.18.5, Transformers 5.14.1, PyTorch 2.9.0
with CUDA 12.8, `bfloat16`, FlashAttention 2.8.3, and batch size 32. The release
processor is configured with a 1,792 visual-token budget. Comparator values
were read from the live
[ViDoRe V3 MTEB leaderboard](https://mteb-leaderboard.hf.space/benchmark/ViDoRe%28v3%29)
on July 22, 2026.
Model encoding runs in `bfloat16`. Before MaxSim scoring, query and document
embeddings are moved to CPU and converted to `float32`. All reported ViDoRe
results use this FP32 scoring path.
| Model | Computer Science | Energy | FinanceEn | FinanceFr | HR | Industrial | Pharmaceuticals | Physics | **Avg. public** |
| :--- | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: |
| **webAI-ColVec1.1-8b (this model)** | 80.14 | 70.09 | **71.89** | **55.10** | 68.49 | **57.41** | 67.88 | 51.29 | **65.29** |
| [VultronRetriever Prime](https://huggingface.co/vultr/VultronRetrieverPrime-Qwen3.5-8B) | 79.81 | **70.26** | 69.01 | 54.51 | 66.82 | **57.41** | **68.19** | **51.73** | 64.72 |
| [webAI-ColVec1-9b](https://huggingface.co/webAI-Official/webAI-ColVec1-9b) | **80.92** | 69.77 | 68.28 | 53.72 | **70.04** | 57.18 | 67.32 | 48.38 | 64.45 |
| **[webAI-ColVec1.1-4b](https://huggingface.co/webAI-Official/webAI-ColVec1.1-4b)** | 80.35 | 69.26 | 69.12 | 53.21 | 67.02 | 56.30 | 67.07 | 51.36 | 64.21 |
| [VultronRetriever Core](https://huggingface.co/vultr/VultronRetrieverCore-Qwen3.5-4.5B) | 79.77 | 69.19 | 68.93 | 52.02 | 66.10 | 56.11 | 67.45 | 50.18 | 63.72 |
| [Nemotron ColEmbed VL 8B V2](https://huggingface.co/nvidia/nemotron-colembed-vl-8b-v2) | 79.29 | 69.82 | 67.29 | 51.54 | 66.32 | 56.03 | 67.19 | 50.84 | 63.54 |
| [webAI-ColVec1-4b](https://huggingface.co/webAI-Official/webAI-ColVec1-4b) | 79.84 | 68.70 | 68.49 | 51.11 | 67.40 | 55.73 | 65.68 | 50.15 | 63.39 |
| [Tomoro ColQwen3 Embed 8B](https://huggingface.co/TomoroAI/tomoro-colqwen3-embed-8b) | 75.35 | 68.41 | 65.08 | 49.10 | 63.98 | 54.41 | 66.36 | 50.13 | 61.60 |
The current MTEB leaderboard entries named `webAI-ColVec1-4b` and
`webAI-ColVec1-9b` refer to the previous ColVec1 release, not these ColVec1.1
checkpoints.
## 💻 Usage
The processor provides the current retrieval API:
- `process_images(images)` prepares one or more document images.
- `process_queries(texts)` prepares one or more natural-language queries.
- `score_retrieval(query_embeddings, document_embeddings)` computes a MaxSim
score matrix with shape `(number_of_queries, number_of_documents)`.
### Prerequisites and attention backends
The public ViDoRe V3 numbers are reproducible with this pinned FlashAttention 2
environment:
```text
Python 3.12
PyTorch 2.9.0 + CUDA 12.8
Transformers 5.14.1
MTEB 2.18.5
Sentence Transformers 5.6.0
FlashAttention 2.8.3
```
The model also supports FlashAttention 3 on Hopper GPUs (H100/H200). See the
[Dao-AILab FlashAttention repository](https://github.com/dao-ailab/flash-attention)
for FlashAttention 3 installation instructions, then select
`flash_attention_3` when loading in a compatible Hopper environment.
FlashAttention 2 and FlashAttention 3 both preserve bidirectional attention.
FlashAttention 3 can improve throughput on Hopper GPUs, but the published
scores use FlashAttention 2; changing kernels can produce small floating-point
differences. Use FlashAttention 2 when reproducing the table.
### Inference code
```python
from io import BytesIO
import requests
import torch
from PIL import Image
from transformers import AutoModel, AutoProcessor
MODEL_ID = "webAI-Official/webAI-ColVec1.1-8b"
DEVICE = "cuda:0" if torch.cuda.is_available() else "cpu"
ATTN_IMPLEMENTATION = (
"flash_attention_2" if torch.cuda.is_available() else "sdpa"
)
# On an H100/H200 with FlashAttention 3 installed, use:
# ATTN_IMPLEMENTATION = "flash_attention_3"
processor = AutoProcessor.from_pretrained(
MODEL_ID,
trust_remote_code=True,
max_num_visual_tokens=1792,
)
model = AutoModel.from_pretrained(
MODEL_ID,
trust_remote_code=True,
dtype=torch.bfloat16,
attn_implementation=ATTN_IMPLEMENTATION,
device_map=DEVICE,
).eval()
queries = [
"When was the United States Declaration of Independence proclaimed?",
"Who printed the edition of Romeo and Juliet?",
]
document_urls = [
"https://upload.wikimedia.org/wikipedia/commons/8/89/US-original-Declaration-1776.jpg",
"https://upload.wikimedia.org/wikipedia/commons/thumb/4/4c/Romeoandjuliet1597.jpg/500px-Romeoandjuliet1597.jpg",
]
def load_image(url: str) -> Image.Image:
response = requests.get(
url,
headers={"User-Agent": "Mozilla/5.0"},
timeout=30,
)
response.raise_for_status()
return Image.open(BytesIO(response.content)).convert("RGB")
device = next(model.parameters()).device
query_inputs = processor.process_queries(queries)
document_inputs = processor.process_images(
[load_image(url) for url in document_urls]
)
query_inputs = {
key: value.to(device) if isinstance(value, torch.Tensor) else value
for key, value in query_inputs.items()
}
document_inputs = {
key: value.to(device) if isinstance(value, torch.Tensor) else value
for key, value in document_inputs.items()
}
with torch.inference_mode():
query_batch = model(**query_inputs)
document_batch = model(**document_inputs)
query_embeddings = [embedding.cpu() for embedding in query_batch]
document_embeddings = [embedding.cpu() for embedding in document_batch]
scores = processor.score_retrieval(
query_embeddings,
document_embeddings,
output_dtype=torch.float32,
)
print(scores)
print("Best document per query:", scores.argmax(dim=1))
```
The processor loads the release's 1,792 visual-token budget by default. To
reduce memory use, pass a lower `max_num_visual_tokens` value to
`AutoProcessor.from_pretrained`; this changes document granularity and may
change retrieval scores.
## ⚖️ Strengths and limitations
### Strengths
- **Performance:** State-of-the-art retrieval performance on the public ViDoRe
V3 tasks, with excellent multimodal document retrieval results.
- **Complex layouts:** Excellent handling of chart-rich PDFs and
domain-specific documents.
- **End-to-end retrieval:** OCR-free retrieval on unseen multimodal documents
without using an intermediate vision-language model to generate summaries.
- **Multilingualism:** Strong performance on non-English document inputs.
### Limitations
- **Storage Cost:** Still larger than single-vector baselines despite the
smaller token dimension.
## License
Model weights are distributed under the
[webAI Non-Commercial License v1.0](https://huggingface.co/webAI-Official/webAI-ColVec1.1-8b/blob/main/LICENSE.md).
See the repository's `NOTICES.md` for upstream attribution.
## 📚 Citation
```bibtex
@misc{webai_colvec1_1_8b,
title = {webAI-ColVec1.1-8b: A Bidirectional Multi-Vector Model for Visual Document Retrieval},
author = {webAI},
year = {2026},
url = {https://huggingface.co/webAI-Official/webAI-ColVec1.1-8b}
}
```