CCCCyx commited on
Commit
ffe00ff
·
verified ·
1 Parent(s): 641be01

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +28 -0
README.md CHANGED
@@ -26,6 +26,10 @@ tags:
26
 
27
  # MOSS-VL-Instruct-0708 W4A16 NF4
28
 
 
 
 
 
29
  This is the Transformers NF4 release of
30
  [MOSS-VL-Instruct-0708](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Instruct-0708).
31
  It supports image and video inference through the standard MOSS-VL offline
@@ -205,3 +209,27 @@ Full inputs, commands and raw results:
205
  - `config.json`: model and bitsandbytes NF4 configuration.
206
  - `generation_config.json`: standard generation settings with BF16 KV cache.
207
  - `modeling_moss_vl.py`: checkpoint-local offline MOSS-VL code.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
26
 
27
  # MOSS-VL-Instruct-0708 W4A16 NF4
28
 
29
+ MOSS-VL is an open vision-language model family from OpenMOSS, supporting image understanding, long-video understanding, and realtime streaming interaction. This repository provides the W4A16 NF4-quantized checkpoint of MOSS-VL-Instruct-0708.
30
+
31
+ **Technical Report**: [https://arxiv.org/pdf/2608.15045](https://arxiv.org/pdf/2608.15045)
32
+
33
  This is the Transformers NF4 release of
34
  [MOSS-VL-Instruct-0708](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Instruct-0708).
35
  It supports image and video inference through the standard MOSS-VL offline
 
209
  - `config.json`: model and bitsandbytes NF4 configuration.
210
  - `generation_config.json`: standard generation settings with BF16 KV cache.
211
  - `modeling_moss_vl.py`: checkpoint-local offline MOSS-VL code.
212
+
213
+ ## Citation
214
+
215
+ ```bibtex
216
+ @misc{mossvl,
217
+ title = {MOSS-VL Technical Report},
218
+ author = {Wang, Pengyu and Tan, Chenkun and Zhou, Shaojun and Zhou, Qirui and Chen, Yanxin and He, Xingyang and Zeng, Huazheng and Cheng, Jijun and Wang, Chenghao and Qian, Xiaomeng and Wang, Pengfei and Huang, Zhan and Gao, Shanqing and Huang, Wei and Cao, Longjun and Ran, Wu and Liu, Jie and Zhu, Changtai and Wang, Hongkai and Tian, Yixian and Liu, Chenghao and Ye, Zhen and Wang, Xinghao and Jiang, Botian and Feng, Guoguo and Fei, Zhaoye and Li, Ruixiao and Chen, Mingshu and Gao, Yang and Cheng, Qinyuan and Li, Shimin and Qiu, Xipeng},
219
+ year = {2026},
220
+ eprint = {2608.15045},
221
+ archivePrefix = {arXiv},
222
+ primaryClass = {cs.CV},
223
+ url = {https://arxiv.org/abs/2608.15045}
224
+ }
225
+
226
+ @misc{mossvideopreview,
227
+ title = {{MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention}},
228
+ author = {Pengyu Wang and Chenkun Tan and Shaojun Zhou and Wei Huang and Qirui Zhou and Zhan Huang and Zhen Ye and Jijun Cheng and Xiaomeng Qian and Yanxin Chen and Xingyang He and Huazheng Zeng and Chenghao Wang and Pengfei Wang and Hongkai Wang and Shanqing Gao and Yixian Tian and Chenghao Liu and Xinghao Wang and Botian Jiang and Xipeng Qiu},
229
+ year = {2026},
230
+ eprint = {2606.07639},
231
+ archivePrefix = {arXiv},
232
+ primaryClass = {cs.CV},
233
+ url = {https://arxiv.org/abs/2606.07639}
234
+ }
235
+ ```