Instructions to use openEuler/smolvla with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use openEuler/smolvla with LeRobot:
# See https://github.com/huggingface/lerobot?tab=readme-ov-file#installation for more details git clone https://github.com/huggingface/lerobot.git cd lerobot pip install -e .[smolvla]
# Launch finetuning on your dataset python lerobot/scripts/train.py \ --policy.path=openEuler/smolvla \ --dataset.repo_id=lerobot/svla_so101_pickplace \ --batch_size=64 \ --steps=20000 \ --output_dir=outputs/train/my_smolvla \ --job_name=my_smolvla_training \ --policy.device=cuda \ --wandb.enable=true
# Run the policy using the record function python -m lerobot.record \ --robot.type=so101_follower \ --robot.port=/dev/ttyACM0 \ # <- Use your port --robot.id=my_blue_follower_arm \ # <- Use your robot id --robot.cameras="{ front: {type: opencv, index_or_path: 8, width: 640, height: 480, fps: 30}}" \ # <- Use your cameras --dataset.single_task="Grasp a lego block and put it in the bin." \ # <- Use the same task description you used in your dataset recording --dataset.repo_id=HF_USER/dataset_name \ # <- This will be the dataset name on HF Hub --dataset.episode_time_s=50 \ --dataset.num_episodes=10 \ --policy.path=openEuler/smolvla - Notebooks
- Google Colab
- Kaggle
Model Card for SmolVLA (IB-Robot)
SmolVLA (Small Vision-Language-Action) policy fine-tuned within the IB-Robot framework. Combines a SmolVLM2-500M vision-language backbone with an action expert for robotic manipulation, packaged with RKNN compiled artifacts for Rockchip RK3588 edge deployment.
Repository Structure
inference_manifest.jsonโ deployment routing (schema v3)config.jsonโ LeRobot policy config (type=smolvla)model.safetensorsโ policy torch weights (~865 MB)policy_preprocessor.json+policy_postprocessor.jsonโ normalization stepsHuggingFaceTB/SmolVLM2-500M-Video-Instruct/โ vendored VLM backbone (12 files, ~1.9 GB)artifacts/rknn/rknn_rk3588/โ RKNN compiled modules (5 artifacts)train_config.jsonโ full training hyperparameters
Deployment Backends
| Target | Backend | Runtime | Hardware |
|---|---|---|---|
rknn_rk3588 |
rknn | rknn-lite2 | Rockchip RK3588 |
torch-cpu |
torch | PyTorch | CPU |
torch-cuda |
torch | PyTorch | NVIDIA GPU |
The RKNN deployment runs a 5-stage pipeline: vision_top / vision_wrist (shared vision encoder) -> embedding -> prefill -> action.
Inputs: observation.state [6], observation.current [6], observation.images.top [3,480,640], observation.images.wrist [3,480,640]
Output: action [6] (5 joints + gripper)
Source Model
This bundle's policy weights are fine-tuned from the upstream SmolVLA base model:
- Policy base model (HuggingFace): lerobot/smolvla_base
- VLM backbone (HuggingFace): HuggingFaceTB/SmolVLM2-500M-Video-Instruct
The VLM backbone is vendored locally under HuggingFaceTB/SmolVLM2-500M-Video-Instruct/ for offline deployment. The RKNN artifacts were converted from the torch weights. See scripts/train_policy.sh for training and scripts/convert_hmm.sh for conversion procedures.
Citation
@inproceedings{smolvla,
title = {SmolVLA: Democratizing Cost-Efficient Vision-Language-Action Models for Robot Manipulation},
author = {LeCun, Yann and others},
booktitle = {HuggingFace},
year = {2025}
}
@software{ib_robot,
title = {IB-Robot: Intelligence Boom Robot},
url = {https://gitcode.com/openeuler/IB_Robot},
license = {Apache-2.0}
}
- Downloads last month
- 14
Model tree for openEuler/smolvla
Base model
lerobot/smolvla_base