ProAR: Learning Prospective Reasoning with Autoregressive Video Models

ProAR solves reasoning tasks by autoregressively generating a sequence of visual states. It equips autoregressive video models with two complementary components:

  • Outcome Guidance predicts a future goal frame to guide current-chunk generation.
  • Transition Guidance aligns the current representation with the clean next-chunk representation during training.

ProAR is built upon Wan2.2-TI2V-5B.

Model Zoo

Dataset Model Checkpoint Training step
VBVR-10Tasks AR vbvr-10tasks/ar/model.pt 10,000
VBVR-10Tasks ProAR vbvr-10tasks/proar/model.pt 10,000
VideoRLVR-3Tasks AR videorlvr-3tasks/ar/model.pt 5,000
VideoRLVR-3Tasks ProAR videorlvr-3tasks/proar/model.pt 5,000
WorldArena AR worldarena/ar/model.pt 10,000
WorldArena ProAR worldarena/proar/model.pt 10,000

Download

Download all released checkpoints:

hf download LinghuiShen/ProAR \
  --local-dir checkpoints/ProAR
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for LinghuiShen/ProAR

Finetuned
(103)
this model