ThinkV2V: Unleashing the Reasoning Capability of MLLMs for Instruction-Guided Video Editing
Abstract
Instruction-guided video editing has made significant progress, yet existing methods use multimodal large language models (MLLMs) primarily as semantic encoders, so they often fall short in working with implicit edits that require causal or semantic reasoning. To bridge this fundamental gap in video editing, we propose ThinkV2V, a reasoning-driven framework for complex instruction-guided video editing, explicitly activating MLLM thinking before visual generation. At its core, ThinkV2V builds on a practical MLLM-to-DiT architecture to turn explicit thinking over the source video and instruction into refined conditioning signals for video editing. Further, we equip it with a dedicated training and inference recipe, combining Progressive Curriculum Training, which gradually cultivates the model from basic editing to reasoning-intensive cases, with Inference-Time Thinking Scaling, which iteratively refines candidate prompts and selects the most reliable one, to better elicit reasoning in challenging editing scenarios. We also curate the ThinkV2V-150K dataset and introduce ThinkV2V-Bench to support training and evaluation of video editing with implicit intent and causal reasoning. Experimental results demonstrate the state-of-the-art performance of ThinkV2V on both complex and standard editing scenarios, in which our 5B-scale DiT model substantially outperforms larger 10B-scale baselines.
Community
- Paper: https://arxiv.org/pdf/2609.38541
- Code: https://github.com/Correr-Zhou/ThinkV2V
- Model (ThinkV2V-5B): https://huggingface.co/donghao-zhou/ThinkV2V-5B
- Dataset #1 (OpenVE-HQ-1M): https://huggingface.co/datasets/donghao-zhou/OpenVE-HQ-1M
- Dataset #2 (ThinkV2V-150K): https://huggingface.co/datasets/donghao-zhou/ThinkV2V-150K
- Benchmark (ThinkV2V-Bench): https://huggingface.co/datasets/donghao-zhou/ThinkV2V-Bench
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing (2026)
- VicEdit: Learning to Edit Videos from Visual In-Context Examples (2026)
- VideoX-Qwen: Data-Centric Instruction-Based Video Editing (2026)
- SenseNova-U1.5: Towards Native Unified Visual Intelligence (2026)
- From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models (2026)
- OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing (2026)
- From Dense Prediction to Visual Editing: Structured Supervision for Unified Image and Video Creation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper