FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos
Abstract
FoldingAgent uses a vision-language model with specialized tools to convert origami videos into executable parametric folding programs via sequential reasoning and physical verification.
We present FoldingAgent, an agentic framework for inferring explicit parametric folding programs directly from origami demonstration videos. Our framework leverages the reasoning power of a pre-trained Vision-Language Model (VLM) equipped with a suite of specialized tools that enable the agent to simulate geometric transitions, verify physical plausibility, retrieve and compare visual content, and evaluate its own predictions. To translate visual content into folding programs, we define a parametric space that consists of the paper's geometry and a set of parametric folding actions. Unlike models that predict static crease patterns, our agent operates sequentially and possesses the ability to re-plan its actions, effectively mitigating the compounding errors inherent in multi-step folding. Our approach takes a step toward closing the gap between human origami knowledge, which is primarily shared through unstructured visual demonstrations, and computational methods, which typically rely on structured, parametric representations such as a crease pattern or an executable parametric plan. We evaluate our approach on PurelandFold, a newly curated benchmark of diverse Pureland origami videos with ground-truth geometry and action labels. Our results demonstrate that by combining VLM reasoning with a set of specialized tools and physical simulation, we can successfully transform unstructured visual demonstrations into executable, physically plausible folding procedures.
Community
This work shows a very cool application of origami representation reconstruction from a natural video. The domain is fun, practical and underexplored. Its use of an agentic framework is practically relevant to the field.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- ArtiMo: Agent-Driven Articulated Mesh Animation (2026)
- RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures (2026)
- FORGE: Towards Functional Tool-Use Generalization via Keypoint Trajectory Reasoning (2026)
- NeoWorld-Pro: Programming Interactive Scenes from Monocular Images for Embodied Simulation (2026)
- aDSL: Agentic 3D Creation via Joint Agent-Program Design (2026)
- Analytic Dynamics: Learning Physics-Grounded Representation for Fast Intrinsic Dynamics Inference from Monocular Videos (2026)
- Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.00377 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper