Image-Text-to-Text
VisCo / README.md
zpatrick's picture
Update README.md
87a6d5e verified
|
Raw
History Blame Contribute Delete
8.3 kB
---
license: apache-2.0
pipeline_tag: image-text-to-text
---
<div align="center">
<h1>🚀 VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression</h1>
<h3>ACM Multimedia 2026</h3>
<p align="center">
<a href="https://github.com/Zyvpeng">Yupeng Zheng</a><sup>*</sup>,
<a href="https://github.com/Jacky-hate">Kai Zou</a><sup>*</sup>,
<a href="https://scholar.google.com/citations?user=kReWULQAAAAJ&hl=en&oi=ao">Bin Liu</a><sup></sup>,
<a href="https://scholar.google.com/citations?user=7620QAMAAAAJ&hl=en&oi=ao">Nenghai Yu</a><sup></sup>
</p>
<p align="center">
<sup>*</sup>Equal contribution &nbsp;&nbsp;
<sup></sup>Corresponding author
</p>
<p align="center">
<a href="https://arxiv.org/abs/2607.12756">
<img src="https://img.shields.io/badge/arXiv-2607.12756-b31b1b.svg">
</a>
<!-- <a href="https://your-project-page.github.io">
<img src="https://img.shields.io/badge/Project-Page-blue">
</a>-->
<a href="https://huggingface.co/zpatrick/VisCo">
<img src="https://img.shields.io/badge/🤗%20Model-HuggingFace-yellow">
</a>
<a href="https://github.com/Zyvpeng/VisCo/blob/main/LICENSE">
<img src="https://img.shields.io/badge/License-Apache--2.0-green">
</a>
<a href="https://github.com/Zyvpeng/VisCo">
<img src="https://img.shields.io/github/stars/Zyvpeng/VisCo?style=social">
</a>
</p>
<p align="center">
<strong>
Compress hundreds of visual tokens into only a few memory tokens while preserving multimodal understanding.
</strong>
</p>
</div>
## 🔥 Highlights
- **Intrinsic self-compression.** VisCo reuses the pretrained VLM itself as a parameter-sharing autoencoder instead of introducing a separate compression network.
- **Lightweight adaptation.** Only memory tokens and encoder-side LoRA adapters are trained; the shared pretrained backbone is kept intact for decoding.
- **Hierarchical KV passing.** Layer-wise key-value states associated with memory tokens are transferred from the encoder directly into the decoder KV cache.
- **Extreme compression.** VisCo retains 85.3% of LLaVA-1.5-7B performance with one visual token and 91.4% of Qwen2-VL-2B performance with 18 tokens.
- **Broad evaluation.** Experiments cover LLaVA-1.5-7B, Qwen2-VL-2B, and Qwen2-VL-7B on six multimodal benchmarks.
## 📊 Main Results
### LLaVA-1.5-7B
Average is the mean percentage of performance retained relative to the original model.
| Method | Tokens | GQA | MMB | MMB-CN | MME | POPE | MMVet | Average |
|:--|--:|--:|--:|--:|--:|--:|--:|--:|
| LLaVA-1.5-7B | 576 | 62.0 | 64.3 | 58.3 | 1510.7 | 85.9 | 31.1 | 100.0% |
| FastV | 32 | 41.5 | 37.8 | 33.2 | 884.6 | 32.5 | 20.7 | 57.6% |
| SparseVLM | 32 | 48.3 | 51.4 | 40.6 | 1046.7 | 67.9 | 18.6 | 72.6% |
| PruMerge+ | 32 | 51.1 | 56.8 | 47.0 | 940.8 | 70.9 | 21.4 | 77.5% |
| DivPrune | 32 | 54.9 | 57.6 | 49.1 | **1284.9** | 81.5 | 26.3 | 87.8% |
| VisPruner | 32 | 52.2 | 58.4 | 52.7 | 1271.0 | 72.7 | **28.8** | 87.8% |
| **VisCo** | **32** | **58.5** | **62.3** | **57.2** | 1152.9 | **81.9** | 27.9 | **91.8%** |
| PruMerge+ | 1 | 26.4 | 13.7 | 13.7 | 568.9 | 40.4 | 12.5 | 35.4% |
| VisPruner | 1 | 41.8 | 22.4 | 25.8 | 764.4 | 49.0 | 12.0 | 48.8% |
| **VisCo** | **1** | **58.2** | **59.0** | **52.4** | **1191.2** | **78.2** | **20.7** | **85.3%** |
### Qwen2-VL
| Base model | Method | Tokens | GQA | MMB | MMB-CN | MME | POPE | MMVet | Average |
|:--|:--|--:|--:|--:|--:|--:|--:|--:|--:|
| Qwen2-VL-2B | Origin | 144 | 59.8 | 67.3 | 61.8 | 1465.8 | 81.5 | 42.5 | 100.0% |
| Qwen2-VL-2B | VisionZip | 36 | 53.2 | 64.3 | 56.7 | 1372.4 | 75.3 | 35.5 | 91.0% |
| Qwen2-VL-2B | **VisCo** | **36** | **61.1** | 64.0 | **58.9** | 1368.3 | **82.9** | **35.6** | **95.2%** |
| Qwen2-VL-2B | VisionZip | 18 | 48.5 | 55.2 | 41.6 | 1300.9 | 62.7 | 25.8 | 76.1% |
| Qwen2-VL-2B | **VisCo** | **18** | **59.7** | **62.9** | **57.8** | **1358.0** | **82.4** | **28.9** | **91.4%** |
| Qwen2-VL-7B | Origin | 144 | 64.8 | 76.1 | 71.6 | 1664.8 | 82.6 | 56.1 | 100.0% |
| Qwen2-VL-7B | VisionZip | 36 | 58.3 | 72.4 | 63.4 | 1530.0 | 75.3 | 39.1 | 87.7% |
| Qwen2-VL-7B | **VisCo** | **36** | **62.6** | **74.0** | **70.0** | **1551.0** | **83.3** | **40.4** | **92.9%** |
| Qwen2-VL-7B | VisionZip | 18 | 55.3 | 66.4 | 58.3 | 1443.8 | 71.7 | 33.9 | 81.3% |
| Qwen2-VL-7B | **VisCo** | **18** | **61.8** | **70.0** | **68.8** | **1496.5** | **81.7** | **39.2** | **90.4%** |
## 🛠️ Installation
```bash
git clone https://github.com/Zyvpeng/VisCo.git
cd VisCo
conda create -n visco python=3.10 -y
conda activate visco
pip install --upgrade pip
pip install torch transformers peft datasets safetensors qwen-vl-utils
pip install flash-attn --no-build-isolation
```
The current implementation additionally uses `torch`, `transformers`, `peft`, `datasets`, `safetensors`, `qwen-vl-utils`, and FlashAttention for LLaVA.
## ⚡ Quick Start
The case examples read `/home/ypzheng/VLM_ICAE/data/case/data.json`. Each item must provide `image` and `question` fields. Run commands from the repository parent so that `VisCo` is importable.
### Qwen2-VL case inference
```bash
cd /home/ypzheng
conda activate mmicae2
python -m VisCo.inference.case_qwen2 \
--model_name_or_path /home/ypzheng/Qwen2-VL-2B-Instruct \
--output_dir /path/to/qwen2_visco_checkpoint.pt \
--fixed_mem_size 18 \
--train false \
--per_device_train_batch_size 1
```
### LLaVA-1.5 case inference
```bash
cd /home/ypzheng
conda activate mmicae2
python -m VisCo.inference.case_llava15 \
--model_name_or_path /home/ypzheng/llava15 \
--output_dir /path/to/llava_visco_checkpoint.pt \
--fixed_mem_size 32 \
--train false \
--per_device_train_batch_size 1
```
`--output_dir` is used by the inference scripts as the checkpoint file path and is loaded with `torch.load`.
## 🚂 Training
Both training scripts fine-tune for one epoch on LLaVA-665K-style JSON data. Update the dataset path near the top of the selected script before launching. Qwen2-VL trains LoRA on `q_proj` and `v_proj`; LLaVA-1.5 trains LoRA on the attention `q/k/v/o` projections. Memory-token embeddings are learned jointly, while decoding reuses the shared backbone without the encoder LoRA adapters.
### Qwen2-VL-2B
```bash
cd /home/ypzheng
torchrun --nproc_per_node=1 -m VisCo.train.train_qwen2 \
--model_name_or_path /home/ypzheng/Qwen2-VL-2B-Instruct \
--output_dir /path/to/output/qwen2_visco \
--fixed_mem_size 18
```
### LLaVA-1.5-7B
```bash
cd /home/ypzheng
torchrun --nproc_per_node=8 -m VisCo.train.train_llava15 \
--model_name_or_path /home/ypzheng/llava15 \
--output_dir /path/to/output/llava_visco \
--fixed_mem_size 32
```
Training uses the tokenization/mapping functions and data collators in `utils/preprocess.py`, then delegates optimization and checkpointing to `train_model` in the same module.
## 🧪 Evaluation
The released evaluation entry points cover MME for both backbones. They load the checkpoint passed through `--output_dir`, preprocess the benchmark with the matching backbone-specific mapping function, and report MME accuracy and accuracy+ scores.
```bash
# Qwen2-VL
python -m VisCo.inference.mme_qwen2 \
--model_name_or_path /home/ypzheng/Qwen2-VL-2B-Instruct \
--output_dir /path/to/qwen2_visco_checkpoint.pt \
--fixed_mem_size 18 \
--train false
# LLaVA-1.5
python -m VisCo.inference.mme_llava15 \
--model_name_or_path /home/ypzheng/llava15 \
--output_dir /path/to/llava_visco_checkpoint.pt \
--fixed_mem_size 32 \
--train false
```
The paper evaluates GQA, MMBench, MMBench-CN, MME, POPE, and MMVet. Additional benchmark entry points will be added as they are cleaned for release.
## 📌 Citation
If you find this project useful, please cite:
```bibtex
@inproceedings{zheng2026visco,
title = {VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression},
author = {Zheng, Yupeng and Zou, Kai and Liu, Bin and Yu, Nenghai},
booktitle = {Proceedings of the 34th ACM International Conference on Multimedia},
year = {2026}
}
```
## 🙏 Acknowledgements
This project builds on [LLaVA](https://github.com/haotian-liu/LLaVA), [Qwen2-VL](https://github.com/QwenLM/Qwen2-VL), [Hugging Face Transformers](https://github.com/huggingface/transformers).