Image-Text-to-Text
RKLLM
English
Chinese
rknn
rk3588
rockchip
npu
quantized
vision-language
multimodal
Instructions to use GatekeeperZA/InternVL3.5-4B-Instruct-RKLLM-v1.2.3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- RKLLM
How to use GatekeeperZA/InternVL3.5-4B-Instruct-RKLLM-v1.2.3 with RKLLM:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
| license: apache-2.0 | |
| base_model: OpenGVLab/InternVL3-4B | |
| tags: | |
| - rkllm | |
| - rknn | |
| - rk3588 | |
| - rockchip | |
| - npu | |
| - quantized | |
| - vision-language | |
| - multimodal | |
| language: | |
| - en | |
| - zh | |
| pipeline_tag: image-text-to-text | |
| # InternVL3.5-4B-Instruct β RKLLM v1.2.3 (w8a8, RK3588) | |
| RKLLM/RKNN conversion of [OpenGVLab/InternVL3-4B](https://huggingface.co/OpenGVLab/InternVL3-4B) for Rockchip RK3588 NPU inference. | |
| Converted with RKLLM Toolkit v1.2.3 (language model) and RKNN Toolkit (vision encoder). This is a multimodal vision-language model β it accepts both images and text as input. | |
| ## Key Details | |
| | Property | Value | | |
| |----------|-------| | |
| | Base Model | OpenGVLab/InternVL3-4B | | |
| | Toolkit Version | RKLLM Toolkit v1.2.3 / RKNN Toolkit | | |
| | Runtime Version | RKLLM Runtime β₯ v1.2.1 + RKNN Runtime | | |
| | Quantization | w8a8 (8-bit weights, 8-bit activations) | | |
| | Target Platform | RK3588 | | |
| | NPU Cores | 3 | | |
| | Thinking Mode | β Not applicable | | |
| | Model Type | Vision-Language (VLM) | | |
| | Languages | English, Chinese (multilingual) | | |
| ## Why This Model? | |
| InternVL3.5-4B is Shanghai AI Lab's compact vision-language model. It provides image understanding, visual question answering, and OCR capabilities at 4B parameters β all running on the RK3588 NPU without a GPU. | |
| Compared to the Qwen3-VL-4B, InternVL has a different training lineage and excels at dense image analysis and chart/document understanding. | |
| ## Hardware Tested | |
| - **Orange Pi 5 Plus** β RK3588, 16GB RAM, Armbian Linux | |
| - RKNPU driver 0.9.8 | |
| - RKLLM Runtime v1.2.3 | |
| ## Usage | |
| ### With the RKLLM API Server (VLM mode) | |
| ```bash | |
| mkdir -p ~/models/internvl3.5-4b | |
| cd ~/models/internvl3.5-4b | |
| git lfs install && git clone https://huggingface.co/GatekeeperZA/InternVL3.5-4B-Instruct-RKLLM-v1.2.3 . | |
| ``` | |
| Use with [GatekeeperZA/RKLLM-API-Server](https://github.com/GatekeeperZA/RKLLM-API-Server) β the server loads both the `.rkllm` and `.rknn` files automatically when placed in the same directory. | |
| ## File Listing | |
| | File | Description | | |
| |------|-------------| | |
| | `internvl3_5-4b-instruct_w8a8_rk3588.rkllm` | Language model weights for RK3588 NPU | | |
| | `internvl3_5-4b_vision_rk3588.rknn` | Vision encoder for RK3588 NPU | | |
| ## Compatibility Notes | |
| - Minimum runtime: RKLLM Runtime v1.2.1 + RKNN Runtime v2.x. v1.2.3 recommended. | |
| - RKNPU driver: β₯ 0.9.6 | |
| - SoCs: RK3588 / RK3588S. Not compatible with RK3576 without reconversion. | |
| - RAM: ~5.5GB loaded. Requires 8GB+ board (16GB recommended). | |
| ## Acknowledgements | |
| - Shanghai AI Lab / OpenGVLab for InternVL3 | |
| - Rockchip / airockchip for the RKLLM and RKNN toolkits | |
| - Converted by [GatekeeperZA](https://huggingface.co/GatekeeperZA) | |