--- library_name: transformers pipeline_tag: image-text-to-text base_model: Qwen/Qwen3-VL-30B-A3B-Instruct tags: - robotics - embodied-ai - video-understanding - progress-estimation - reward-modeling - qwen3-vl license: other --- # VLAC-Cut: Video Progress Estimation for Process-Level Robot Rollout Segmentation
[Paper](https://arxiv.org/abs/2607.09776) · [Code](https://github.com/InternRobotics/VLAC-cut) · [Model](https://huggingface.co/InternRobotics/VLAC-Cut) · [Benchmark](https://huggingface.co/datasets/InternRobotics/VLAC-Cut-Benchmark)
## Overview **VLAC-Cut** is a process-level multimodal trajectory critic for robot post-training data curation. Given a natural-language task instruction, an optional task plan, and a robot rollout video, VLAC-Cut estimates signed task progress over time and identifies temporal segments associated with task advancement or degradation. Unlike methods that assume task progress increases monotonically over time, VLAC-Cut models non-monotonic execution dynamics, including advancement, stagnation, regression, and recovery. This formulation supports process-level analysis of partial completion, temporary failure, subsequent recovery, and rollout segmentation for post-training data selection. This Hugging Face repository contains the VLAC-Cut model weights and loading assets. The official inference examples and evaluation code are maintained in the GitHub repository. ## Highlights * **Video-level temporal reasoning:** Analyzes robot execution videos rather than isolated images or image pairs and identifies temporal segments associated with task advancement or degradation. * **Non-monotonic progress estimation:** Captures advancement, stagnation, regression, and recovery without imposing a monotonically increasing progress assumption. * **Zero-shot generalization:** Generalizes across manipulation tasks, scenes, object configurations, and camera viewpoints. * **Flexible temporal resolution:** Supports configurable video sampling frequencies for both coarse- and fine-grained progress estimation. ## Model Overview | Property | Description | |---|---| | Base model | `Qwen/Qwen3-VL-30B-A3B-Instruct` | | Input | Task instruction, optional task plan, and sampled video frames | | Output | Timestamped task-progress estimates | | Sampling rate | `2 Hz`-`20 Hz` | | Default sampling rate | `2.0 Hz` | ## Load with Transformers ```python from transformers import AutoModelForImageTextToText, AutoProcessor model_id = "InternRobotics/VLAC-Cut" processor = AutoProcessor.from_pretrained(model_id) model = AutoModelForImageTextToText.from_pretrained( model_id, dtype="auto", device_map="auto", ) model.eval() ``` ## Quick Start Run progress inference on a local video using the GitHub code: ```bash git clone https://github.com/InternRobotics/VLAC-cut cd VLAC-cut python scripts/run_example.py \ --model-path InternRobotics/VLAC-Cut \ --video-path \ --task-instruction "" \ --task-plan $'' \ --output-jsonl ``` Render a prediction JSONL file as an annotated video: ```bash python scripts/utils/render_prediction_video.py \ --input-jsonl \ --output-video ``` ## Citation Please cite the following paper when using VLAC-Cut, the released model, or the Video Progress Benchmark: ```bibtex @misc{zhai2026helphumanefficientlargescalerobot, title={HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation}, author={Shaopeng Zhai and Qi Zhang and Tianyi Zhang and Haoran Zhang and Fuxian Huang and Zhanhui Lin and Zijun Xu and Weinan Zhang}, year={2026}, eprint={2607.09776}, archivePrefix={arXiv}, primaryClass={cs.RO}, url={https://arxiv.org/abs/2607.09776}, } ``` ## License The model weights and third-party training data may be subject to additional licenses or terms of use. The source code in the GitHub repository is released under the MIT License.