| --- |
| license: apache-2.0 |
| pipeline_tag: image-text-to-text |
| --- |
| <div align="center"> |
|
|
| <h1>🚀 VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression</h1> |
|
|
| <h3>ACM Multimedia 2026</h3> |
|
|
| <p align="center"> |
| <a href="https://github.com/Zyvpeng">Yupeng Zheng</a><sup>*</sup>, |
| <a href="https://github.com/Jacky-hate">Kai Zou</a><sup>*</sup>, |
| <a href="https://scholar.google.com/citations?user=kReWULQAAAAJ&hl=en&oi=ao">Bin Liu</a><sup>†</sup>, |
| <a href="https://scholar.google.com/citations?user=7620QAMAAAAJ&hl=en&oi=ao">Nenghai Yu</a><sup>†</sup> |
| </p> |
|
|
| <p align="center"> |
| <sup>*</sup>Equal contribution |
| <sup>†</sup>Corresponding author |
| </p> |
| |
| <p align="center"> |
| <a href="https://arxiv.org/abs/2607.12756"> |
| <img src="https://img.shields.io/badge/arXiv-2607.12756-b31b1b.svg"> |
| </a> |
| <!-- <a href="https://your-project-page.github.io"> |
| <img src="https://img.shields.io/badge/Project-Page-blue"> |
| </a>--> |
| <a href="https://huggingface.co/zpatrick/VisCo"> |
| <img src="https://img.shields.io/badge/🤗%20Model-HuggingFace-yellow"> |
| </a> |
| <a href="https://github.com/Zyvpeng/VisCo/blob/main/LICENSE"> |
| <img src="https://img.shields.io/badge/License-Apache--2.0-green"> |
| </a> |
| <a href="https://github.com/Zyvpeng/VisCo"> |
| <img src="https://img.shields.io/github/stars/Zyvpeng/VisCo?style=social"> |
| </a> |
| </p> |
| |
| <p align="center"> |
| <strong> |
| Compress hundreds of visual tokens into only a few memory tokens while preserving multimodal understanding. |
| </strong> |
| </p> |
| |
| </div> |
| |
| ## 🔥 Highlights |
| |
| - **Intrinsic self-compression.** VisCo reuses the pretrained VLM itself as a parameter-sharing autoencoder instead of introducing a separate compression network. |
| - **Lightweight adaptation.** Only memory tokens and encoder-side LoRA adapters are trained; the shared pretrained backbone is kept intact for decoding. |
| - **Hierarchical KV passing.** Layer-wise key-value states associated with memory tokens are transferred from the encoder directly into the decoder KV cache. |
| - **Extreme compression.** VisCo retains 85.3% of LLaVA-1.5-7B performance with one visual token and 91.4% of Qwen2-VL-2B performance with 18 tokens. |
| - **Broad evaluation.** Experiments cover LLaVA-1.5-7B, Qwen2-VL-2B, and Qwen2-VL-7B on six multimodal benchmarks. |
| |
| |
| ## 📊 Main Results |
| |
| ### LLaVA-1.5-7B |
| |
| Average is the mean percentage of performance retained relative to the original model. |
| |
| | Method | Tokens | GQA | MMB | MMB-CN | MME | POPE | MMVet | Average | |
| |:--|--:|--:|--:|--:|--:|--:|--:|--:| |
| | LLaVA-1.5-7B | 576 | 62.0 | 64.3 | 58.3 | 1510.7 | 85.9 | 31.1 | 100.0% | |
| | FastV | 32 | 41.5 | 37.8 | 33.2 | 884.6 | 32.5 | 20.7 | 57.6% | |
| | SparseVLM | 32 | 48.3 | 51.4 | 40.6 | 1046.7 | 67.9 | 18.6 | 72.6% | |
| | PruMerge+ | 32 | 51.1 | 56.8 | 47.0 | 940.8 | 70.9 | 21.4 | 77.5% | |
| | DivPrune | 32 | 54.9 | 57.6 | 49.1 | **1284.9** | 81.5 | 26.3 | 87.8% | |
| | VisPruner | 32 | 52.2 | 58.4 | 52.7 | 1271.0 | 72.7 | **28.8** | 87.8% | |
| | **VisCo** | **32** | **58.5** | **62.3** | **57.2** | 1152.9 | **81.9** | 27.9 | **91.8%** | |
| | PruMerge+ | 1 | 26.4 | 13.7 | 13.7 | 568.9 | 40.4 | 12.5 | 35.4% | |
| | VisPruner | 1 | 41.8 | 22.4 | 25.8 | 764.4 | 49.0 | 12.0 | 48.8% | |
| | **VisCo** | **1** | **58.2** | **59.0** | **52.4** | **1191.2** | **78.2** | **20.7** | **85.3%** | |
| |
| ### Qwen2-VL |
| |
| | Base model | Method | Tokens | GQA | MMB | MMB-CN | MME | POPE | MMVet | Average | |
| |:--|:--|--:|--:|--:|--:|--:|--:|--:|--:| |
| | Qwen2-VL-2B | Origin | 144 | 59.8 | 67.3 | 61.8 | 1465.8 | 81.5 | 42.5 | 100.0% | |
| | Qwen2-VL-2B | VisionZip | 36 | 53.2 | 64.3 | 56.7 | 1372.4 | 75.3 | 35.5 | 91.0% | |
| | Qwen2-VL-2B | **VisCo** | **36** | **61.1** | 64.0 | **58.9** | 1368.3 | **82.9** | **35.6** | **95.2%** | |
| | Qwen2-VL-2B | VisionZip | 18 | 48.5 | 55.2 | 41.6 | 1300.9 | 62.7 | 25.8 | 76.1% | |
| | Qwen2-VL-2B | **VisCo** | **18** | **59.7** | **62.9** | **57.8** | **1358.0** | **82.4** | **28.9** | **91.4%** | |
| | Qwen2-VL-7B | Origin | 144 | 64.8 | 76.1 | 71.6 | 1664.8 | 82.6 | 56.1 | 100.0% | |
| | Qwen2-VL-7B | VisionZip | 36 | 58.3 | 72.4 | 63.4 | 1530.0 | 75.3 | 39.1 | 87.7% | |
| | Qwen2-VL-7B | **VisCo** | **36** | **62.6** | **74.0** | **70.0** | **1551.0** | **83.3** | **40.4** | **92.9%** | |
| | Qwen2-VL-7B | VisionZip | 18 | 55.3 | 66.4 | 58.3 | 1443.8 | 71.7 | 33.9 | 81.3% | |
| | Qwen2-VL-7B | **VisCo** | **18** | **61.8** | **70.0** | **68.8** | **1496.5** | **81.7** | **39.2** | **90.4%** | |
| |
| ## 🛠️ Installation |
| |
| ```bash |
| git clone https://github.com/Zyvpeng/VisCo.git |
| cd VisCo |
| |
| conda create -n visco python=3.10 -y |
| conda activate visco |
| pip install --upgrade pip |
| pip install torch transformers peft datasets safetensors qwen-vl-utils |
| pip install flash-attn --no-build-isolation |
| ``` |
| |
| The current implementation additionally uses `torch`, `transformers`, `peft`, `datasets`, `safetensors`, `qwen-vl-utils`, and FlashAttention for LLaVA. |
| |
| ## ⚡ Quick Start |
| |
| The case examples read `/home/ypzheng/VLM_ICAE/data/case/data.json`. Each item must provide `image` and `question` fields. Run commands from the repository parent so that `VisCo` is importable. |
| |
| ### Qwen2-VL case inference |
| |
| ```bash |
| cd /home/ypzheng |
| conda activate mmicae2 |
| |
| python -m VisCo.inference.case_qwen2 \ |
| --model_name_or_path /home/ypzheng/Qwen2-VL-2B-Instruct \ |
| --output_dir /path/to/qwen2_visco_checkpoint.pt \ |
| --fixed_mem_size 18 \ |
| --train false \ |
| --per_device_train_batch_size 1 |
| ``` |
| |
| ### LLaVA-1.5 case inference |
| |
| ```bash |
| cd /home/ypzheng |
| conda activate mmicae2 |
| |
| python -m VisCo.inference.case_llava15 \ |
| --model_name_or_path /home/ypzheng/llava15 \ |
| --output_dir /path/to/llava_visco_checkpoint.pt \ |
| --fixed_mem_size 32 \ |
| --train false \ |
| --per_device_train_batch_size 1 |
| ``` |
| |
| `--output_dir` is used by the inference scripts as the checkpoint file path and is loaded with `torch.load`. |
| |
| ## 🚂 Training |
| |
| Both training scripts fine-tune for one epoch on LLaVA-665K-style JSON data. Update the dataset path near the top of the selected script before launching. Qwen2-VL trains LoRA on `q_proj` and `v_proj`; LLaVA-1.5 trains LoRA on the attention `q/k/v/o` projections. Memory-token embeddings are learned jointly, while decoding reuses the shared backbone without the encoder LoRA adapters. |
| |
| ### Qwen2-VL-2B |
| |
| ```bash |
| cd /home/ypzheng |
| torchrun --nproc_per_node=1 -m VisCo.train.train_qwen2 \ |
| --model_name_or_path /home/ypzheng/Qwen2-VL-2B-Instruct \ |
| --output_dir /path/to/output/qwen2_visco \ |
| --fixed_mem_size 18 |
| ``` |
| |
| ### LLaVA-1.5-7B |
| |
| ```bash |
| cd /home/ypzheng |
| torchrun --nproc_per_node=8 -m VisCo.train.train_llava15 \ |
| --model_name_or_path /home/ypzheng/llava15 \ |
| --output_dir /path/to/output/llava_visco \ |
| --fixed_mem_size 32 |
| ``` |
| |
| Training uses the tokenization/mapping functions and data collators in `utils/preprocess.py`, then delegates optimization and checkpointing to `train_model` in the same module. |
| |
| ## 🧪 Evaluation |
| |
| The released evaluation entry points cover MME for both backbones. They load the checkpoint passed through `--output_dir`, preprocess the benchmark with the matching backbone-specific mapping function, and report MME accuracy and accuracy+ scores. |
| |
| ```bash |
| # Qwen2-VL |
| python -m VisCo.inference.mme_qwen2 \ |
| --model_name_or_path /home/ypzheng/Qwen2-VL-2B-Instruct \ |
| --output_dir /path/to/qwen2_visco_checkpoint.pt \ |
| --fixed_mem_size 18 \ |
| --train false |
| |
| # LLaVA-1.5 |
| python -m VisCo.inference.mme_llava15 \ |
| --model_name_or_path /home/ypzheng/llava15 \ |
| --output_dir /path/to/llava_visco_checkpoint.pt \ |
| --fixed_mem_size 32 \ |
| --train false |
| ``` |
| |
| The paper evaluates GQA, MMBench, MMBench-CN, MME, POPE, and MMVet. Additional benchmark entry points will be added as they are cleaned for release. |
| |
| |
| ## 📌 Citation |
| |
| If you find this project useful, please cite: |
| |
| ```bibtex |
| @inproceedings{zheng2026visco, |
| title = {VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression}, |
| author = {Zheng, Yupeng and Zou, Kai and Liu, Bin and Yu, Nenghai}, |
| booktitle = {Proceedings of the 34th ACM International Conference on Multimedia}, |
| year = {2026} |
| } |
| ``` |
| |
| ## 🙏 Acknowledgements |
| |
| This project builds on [LLaVA](https://github.com/haotian-liu/LLaVA), [Qwen2-VL](https://github.com/QwenLM/Qwen2-VL), [Hugging Face Transformers](https://github.com/huggingface/transformers). |
| |
| |