Papers
arxiv:2609.02886

SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models

Published on Sep 2
· Submitted by
junchao-cuhk
on Sep 3
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,

Abstract

SolarWM provides an open framework and unified training recipe for building interactive video world models across diverse data sources and generator backbones, enabling long-horizon real-time rollouts.

We introduce SolarWM, a fully open foundation for building interactive video world models from data preparation through long-horizon inference. Training across heterogeneous data sources and video backbones is challenging: datasets differ in temporal scale, camera geometry, visual quality, motion, and captioning styles, while video generators use distinct representations and architectures. Naive data mixing and model-specific implementations therefore produce inconsistent supervision and make results difficult to reproduce and compare. SolarWM addresses this coupling with a reconfigurable multi-source data engine and a backbone-native adaptation framework. The engine converts 1.43 million canonical clips from 10 datasets into a unified, frame-aligned contract covering visual observations, metric camera geometry, captions, quality metadata, selection decisions, and provenance, while decoupling source processing from mixture construction. Under shared camera-conditioning, training, and inference interfaces, we instantiate four 5B--33B models based on Wan2.2, LTX-2.5, and MiniMax-H3 while preserving their native representations and objectives. A unified three-stage recipe combines bidirectional adaptation, teacher-forced autoregressive initialization, and distribution matching distillation. The resulting causal models enable real-time interaction over rollouts ranging from minutes to hours after being trained on only 5s sequences. By releasing the resulting data, pipeline, recipes, weights, and framework, SolarWM provides a reproducible and extensible foundation for interactive world-model research.

Community

Paper submitter

We present SolarWM, a fully open foundation for building interactive video
world models from data preparation through scalable training and long-horizon
inference.

  • Open, reconfigurable data infrastructure. SolarWM converts 1.43 million
    canonical clips from 14 datasets into a unified, frame-aligned contract for
    observations, metric camera geometry, captions, quality metadata, selection,
    and provenance. Source processing is decoupled from training-mixture design.
  • A scalable, backbone-native model family. One framework supports four
    5B–33B models across Wan2.2, LTX-2.5, and MiniMax-H3 while preserving each
    backbone's native representation and objective.
  • A simple three-stage training recipe. Bidirectional adaptation,
    teacher-forced autoregressive initialization, and distribution matching
    distillation form a shared route across heterogeneous video backbones,
    without specialized ODE or consistency-distillation initialization.
  • Long-horizon interaction from short training clips. After training only
    on 5-second sequences, the resulting causal models support real-time
    interaction with rollouts spanning minutes to hours, without long-sequence
    fine-tuning or attention-sink mechanisms.

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.02886
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 5

Browse 5 models citing this paper

Datasets citing this paper 1

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.02886 in a Space README.md to link it from this page.

Collections including this paper 4