Papers
arxiv:2609.13053

Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model

Published on Sep 11
· Submitted by
Hoeun Lee
on Sep 15
Authors:
,
,
,
,
,

Abstract

A shared trajectory model unifies goal and dynamics prediction with action generation for language-conditioned robot control, improving adaptation and success through joint denoising and test-time scaling.

Visual goal and dynamics prediction can provide language-conditioned robot policies with both a target outcome and a representation of action-dependent scene changes. We bring these predictions into action generation and selection through a shared trajectory model. Dynin-Robotics implements this formulation on Dynin-Omni, an omnimodal masked-diffusion backbone, representing language, visual observations, goals, and actions as discrete tokens. By varying conditioning and target spans, the same model learns action prediction, action-conditioned next-observation prediction, terminal goal-state prediction, and trajectory-to-instruction reconstruction. These interfaces support test-time scaling through goal prediction, action-candidate evaluation, and joint refinement of action and future-state predictions. We continually pretrain the model on approximately 1.33 million trajectories from 48 Open X-Embodiment datasets and adapt it separately to downstream domains. On two VLABench tasks, robot pretraining improves adaptation within a fixed Stage-2 step budget, and the full objective mixture improves shifted-instruction success over Policy-only post-training under the same coupled decoder. Combining goal guidance with joint action-next-state denoising further improves shifted-instruction success over action-only decoding; the benefit depends on how the predictions are composed. Dynin-Robotics achieves competitive performance on LIBERO and zero-shot LIBERO-Plus, together with a 78.4% average success rate across four manipulation conditions on a Franka Research 3 robot. An optimized block-parallel implementation accelerates model-side action decoding by up to 29.2x relative to the base implementation under the reported profiling setup. These results support shared trajectory modeling as a common interface for learning complementary robot objectives and composing their predictions during control.

Community

Paper author Paper submitter

We introduce Dynin-Robotics, an omnimodal masked-diffusion vision-language-action foundation model that unifies policy generation, world modeling, goal-state prediction, and task understanding within a single architecture. By leveraging shared trajectory modeling and iterative bidirectional refinement, Dynin-Robotics integrates visual predictions into action generation and selection, enabling goal-guided control, joint action–future-state denoising, and compositional test-time scaling. As demonstrated in our experiments, Dynin-Robotics achieves competitive performance across simulation benchmarks and real-world manipulation tasks, supporting discrete diffusion as a practical paradigm for unified robotic understanding, prediction, and control.

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.13053
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.13053 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.13053 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.