Visual Document Retrieval
Transformers
Safetensors
multilingual
qwen3_5
feature-extraction
text
image
multimodal-embedding
vidore
colbert
colqwen3_5
multilingual-embedding
custom_code
Instructions to use webAI-Official/webAI-ColVec1.1-8b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use webAI-Official/webAI-ColVec1.1-8b with Transformers:
# Load model directly from transformers import AutoProcessor, AutoModel processor = AutoProcessor.from_pretrained("webAI-Official/webAI-ColVec1.1-8b", trust_remote_code=True) model = AutoModel.from_pretrained("webAI-Official/webAI-ColVec1.1-8b", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
| pipeline_tag: visual-document-retrieval | |
| library_name: transformers | |
| language: | |
| - multilingual | |
| license: other | |
| license_name: webai-non-commercial-license-v1.0 | |
| license_link: https://huggingface.co/webAI-Official/webAI-ColVec1.1-8b/blob/main/LICENSE.md | |
| base_model: Qwen/Qwen3.5-9B | |
| datasets: | |
| - vidore/colpali_train_set | |
| - Tevatron/docmatix-ir | |
| - openbmb/VisRAG-Ret-Train-In-domain-data | |
| - openbmb/VisRAG-Ret-Train-Synthetic-data | |
| - llamaindex/vdr-multilingual-train | |
| - Tevatron/wiki-ss-nq | |
| tags: | |
| - text | |
| - image | |
| - multimodal-embedding | |
| - visual-document-retrieval | |
| - vidore | |
| - colbert | |
| - colqwen3_5 | |
| - multilingual-embedding | |
| # webAI-Official/webAI-ColVec1.1-8b | |
| ## ⚡ Summary | |
| **webAI-Official/webAI-ColVec1.1-8b** is a ColBERT-style multimodal | |
| embedding model based on | |
| [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B). It maps text | |
| queries and visual documents (images or rendered PDF pages) into aligned, | |
| L2-normalized multi-vector embeddings for late-interaction retrieval. | |
| The model uses bidirectional attention in Qwen3.5's full-attention layers and | |
| a learned 640-dimensional projection head. The unused language-model head has | |
| been removed from the released checkpoint. Consequently, the released | |
| embedding model contains ~8.4B parameters, reflected by the | |
| `webAI-ColVec1.1-8b` repository name, while its original backbone is | |
| Qwen3.5-9B. | |
| ### Training data | |
| We created filtered, balanced, and multilingual curated subsets from six public | |
| datasets: [ColPali Train Set](https://huggingface.co/datasets/vidore/colpali_train_set), | |
| [Docmatix-IR](https://huggingface.co/datasets/Tevatron/docmatix-ir), | |
| [VisRAG In-Domain](https://huggingface.co/datasets/openbmb/VisRAG-Ret-Train-In-domain-data), | |
| [VisRAG Synthetic](https://huggingface.co/datasets/openbmb/VisRAG-Ret-Train-Synthetic-data), | |
| [VDR Multilingual Train](https://huggingface.co/datasets/llamaindex/vdr-multilingual-train), | |
| and [Wiki-SS-NQ](https://huggingface.co/datasets/Tevatron/wiki-ss-nq). This | |
| Qwen3.5-9B-backbone model was trained on a 750,000-sample curated subset as well | |
| as synthetically generated data. | |
| ## 🛠️ Model specifications | |
| | Feature | Detail | | |
| | :--- | :--- | | |
| | **Architecture** | Qwen3.5-9B vision-language model + 640-dimensional linear projection | | |
| | **Released parameters** | 8,395,317,104 | | |
| | **Method** | ColBERT-style late interaction with MaxSim scoring | | |
| | **Output** | L2-normalized multi-vector embeddings `(sequence_length, 640)` | | |
| | **Modalities** | Text queries and document images | | |
| | **Attention** | Bidirectional full-attention layers; selectable FlashAttention 2, FlashAttention 3, or SDPA kernel | | |
| | **Visual-token budget** | 1,792 tokens per image in the released processor | | |
| | **Training** | LoRA adapters and a fully trained projection layer, merged for release | | |
| | **Weights** | `bfloat16`; language-model head removed | | |
| ### Key properties | |
| - **Unified encoder:** The same model encodes text and document images. | |
| - **Token-level retrieval:** Multi-vector embeddings preserve fine-grained | |
| layout and content signals that single-vector pooling can discard. | |
| - **Compact projection:** Hidden states are projected to 640 dimensions | |
| without an activation function. | |
| - **Bidirectional retrieval attention:** Selecting FlashAttention 2 or 3 | |
| changes the execution kernel, not the model's bidirectional attention mode. | |
| ## 📊 Evaluation results | |
| The table reports NDCG@10 scores on the eight **public** ViDoRe V3 tasks as | |
| percentages rather than values between 0 and 1 (for example, 0.80 is shown as | |
| 80.00). Each task value is the mean of its six language subsets; the public average is the unweighted mean of the eight public task values. | |
| The result artifacts record MTEB 2.18.5, Transformers 5.14.1, PyTorch 2.9.0 | |
| with CUDA 12.8, `bfloat16`, FlashAttention 2.8.3, and batch size 32. The release | |
| processor is configured with a 1,792 visual-token budget. Comparator values | |
| were read from the live | |
| [ViDoRe V3 MTEB leaderboard](https://mteb-leaderboard.hf.space/benchmark/ViDoRe%28v3%29) | |
| on July 22, 2026. | |
| Model encoding runs in `bfloat16`. Before MaxSim scoring, query and document | |
| embeddings are moved to CPU and converted to `float32`. All reported ViDoRe | |
| results use this FP32 scoring path. | |
| | Model | Computer Science | Energy | FinanceEn | FinanceFr | HR | Industrial | Pharmaceuticals | Physics | **Avg. public** | | |
| | :--- | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: | | |
| | **webAI-ColVec1.1-8b (this model)** | 80.14 | 70.09 | **71.89** | **55.10** | 68.49 | **57.41** | 67.88 | 51.29 | **65.29** | | |
| | [VultronRetriever Prime](https://huggingface.co/vultr/VultronRetrieverPrime-Qwen3.5-8B) | 79.81 | **70.26** | 69.01 | 54.51 | 66.82 | **57.41** | **68.19** | **51.73** | 64.72 | | |
| | [webAI-ColVec1-9b](https://huggingface.co/webAI-Official/webAI-ColVec1-9b) | **80.92** | 69.77 | 68.28 | 53.72 | **70.04** | 57.18 | 67.32 | 48.38 | 64.45 | | |
| | **[webAI-ColVec1.1-4b](https://huggingface.co/webAI-Official/webAI-ColVec1.1-4b)** | 80.35 | 69.26 | 69.12 | 53.21 | 67.02 | 56.30 | 67.07 | 51.36 | 64.21 | | |
| | [VultronRetriever Core](https://huggingface.co/vultr/VultronRetrieverCore-Qwen3.5-4.5B) | 79.77 | 69.19 | 68.93 | 52.02 | 66.10 | 56.11 | 67.45 | 50.18 | 63.72 | | |
| | [Nemotron ColEmbed VL 8B V2](https://huggingface.co/nvidia/nemotron-colembed-vl-8b-v2) | 79.29 | 69.82 | 67.29 | 51.54 | 66.32 | 56.03 | 67.19 | 50.84 | 63.54 | | |
| | [webAI-ColVec1-4b](https://huggingface.co/webAI-Official/webAI-ColVec1-4b) | 79.84 | 68.70 | 68.49 | 51.11 | 67.40 | 55.73 | 65.68 | 50.15 | 63.39 | | |
| | [Tomoro ColQwen3 Embed 8B](https://huggingface.co/TomoroAI/tomoro-colqwen3-embed-8b) | 75.35 | 68.41 | 65.08 | 49.10 | 63.98 | 54.41 | 66.36 | 50.13 | 61.60 | | |
| The current MTEB leaderboard entries named `webAI-ColVec1-4b` and | |
| `webAI-ColVec1-9b` refer to the previous ColVec1 release, not these ColVec1.1 | |
| checkpoints. | |
| ## 💻 Usage | |
| The processor provides the current retrieval API: | |
| - `process_images(images)` prepares one or more document images. | |
| - `process_queries(texts)` prepares one or more natural-language queries. | |
| - `score_retrieval(query_embeddings, document_embeddings)` computes a MaxSim | |
| score matrix with shape `(number_of_queries, number_of_documents)`. | |
| ### Prerequisites and attention backends | |
| The public ViDoRe V3 numbers are reproducible with this pinned FlashAttention 2 | |
| environment: | |
| ```text | |
| Python 3.12 | |
| PyTorch 2.9.0 + CUDA 12.8 | |
| Transformers 5.14.1 | |
| MTEB 2.18.5 | |
| Sentence Transformers 5.6.0 | |
| FlashAttention 2.8.3 | |
| ``` | |
| The model also supports FlashAttention 3 on Hopper GPUs (H100/H200). See the | |
| [Dao-AILab FlashAttention repository](https://github.com/dao-ailab/flash-attention) | |
| for FlashAttention 3 installation instructions, then select | |
| `flash_attention_3` when loading in a compatible Hopper environment. | |
| FlashAttention 2 and FlashAttention 3 both preserve bidirectional attention. | |
| FlashAttention 3 can improve throughput on Hopper GPUs, but the published | |
| scores use FlashAttention 2; changing kernels can produce small floating-point | |
| differences. Use FlashAttention 2 when reproducing the table. | |
| ### Inference code | |
| ```python | |
| from io import BytesIO | |
| import requests | |
| import torch | |
| from PIL import Image | |
| from transformers import AutoModel, AutoProcessor | |
| MODEL_ID = "webAI-Official/webAI-ColVec1.1-8b" | |
| DEVICE = "cuda:0" if torch.cuda.is_available() else "cpu" | |
| ATTN_IMPLEMENTATION = ( | |
| "flash_attention_2" if torch.cuda.is_available() else "sdpa" | |
| ) | |
| # On an H100/H200 with FlashAttention 3 installed, use: | |
| # ATTN_IMPLEMENTATION = "flash_attention_3" | |
| processor = AutoProcessor.from_pretrained( | |
| MODEL_ID, | |
| trust_remote_code=True, | |
| max_num_visual_tokens=1792, | |
| ) | |
| model = AutoModel.from_pretrained( | |
| MODEL_ID, | |
| trust_remote_code=True, | |
| dtype=torch.bfloat16, | |
| attn_implementation=ATTN_IMPLEMENTATION, | |
| device_map=DEVICE, | |
| ).eval() | |
| queries = [ | |
| "When was the United States Declaration of Independence proclaimed?", | |
| "Who printed the edition of Romeo and Juliet?", | |
| ] | |
| document_urls = [ | |
| "https://upload.wikimedia.org/wikipedia/commons/8/89/US-original-Declaration-1776.jpg", | |
| "https://upload.wikimedia.org/wikipedia/commons/thumb/4/4c/Romeoandjuliet1597.jpg/500px-Romeoandjuliet1597.jpg", | |
| ] | |
| def load_image(url: str) -> Image.Image: | |
| response = requests.get( | |
| url, | |
| headers={"User-Agent": "Mozilla/5.0"}, | |
| timeout=30, | |
| ) | |
| response.raise_for_status() | |
| return Image.open(BytesIO(response.content)).convert("RGB") | |
| device = next(model.parameters()).device | |
| query_inputs = processor.process_queries(queries) | |
| document_inputs = processor.process_images( | |
| [load_image(url) for url in document_urls] | |
| ) | |
| query_inputs = { | |
| key: value.to(device) if isinstance(value, torch.Tensor) else value | |
| for key, value in query_inputs.items() | |
| } | |
| document_inputs = { | |
| key: value.to(device) if isinstance(value, torch.Tensor) else value | |
| for key, value in document_inputs.items() | |
| } | |
| with torch.inference_mode(): | |
| query_batch = model(**query_inputs) | |
| document_batch = model(**document_inputs) | |
| query_embeddings = [embedding.cpu() for embedding in query_batch] | |
| document_embeddings = [embedding.cpu() for embedding in document_batch] | |
| scores = processor.score_retrieval( | |
| query_embeddings, | |
| document_embeddings, | |
| output_dtype=torch.float32, | |
| ) | |
| print(scores) | |
| print("Best document per query:", scores.argmax(dim=1)) | |
| ``` | |
| The processor loads the release's 1,792 visual-token budget by default. To | |
| reduce memory use, pass a lower `max_num_visual_tokens` value to | |
| `AutoProcessor.from_pretrained`; this changes document granularity and may | |
| change retrieval scores. | |
| ## ⚖️ Strengths and limitations | |
| ### Strengths | |
| - **Performance:** State-of-the-art retrieval performance on the public ViDoRe | |
| V3 tasks, with excellent multimodal document retrieval results. | |
| - **Complex layouts:** Excellent handling of chart-rich PDFs and | |
| domain-specific documents. | |
| - **End-to-end retrieval:** OCR-free retrieval on unseen multimodal documents | |
| without using an intermediate vision-language model to generate summaries. | |
| - **Multilingualism:** Strong performance on non-English document inputs. | |
| ### Limitations | |
| - **Storage Cost:** Still larger than single-vector baselines despite the | |
| smaller token dimension. | |
| ## License | |
| Model weights are distributed under the | |
| [webAI Non-Commercial License v1.0](https://huggingface.co/webAI-Official/webAI-ColVec1.1-8b/blob/main/LICENSE.md). | |
| See the repository's `NOTICES.md` for upstream attribution. | |
| ## 📚 Citation | |
| ```bibtex | |
| @misc{webai_colvec1_1_8b, | |
| title = {webAI-ColVec1.1-8b: A Bidirectional Multi-Vector Model for Visual Document Retrieval}, | |
| author = {webAI}, | |
| year = {2026}, | |
| url = {https://huggingface.co/webAI-Official/webAI-ColVec1.1-8b} | |
| } | |
| ``` | |