ProAR: Learning Prospective Reasoning with Autoregressive Video Models
ProAR solves reasoning tasks by autoregressively generating a sequence of visual states. It equips autoregressive video models with two complementary components:
- Outcome Guidance predicts a future goal frame to guide current-chunk generation.
- Transition Guidance aligns the current representation with the clean next-chunk representation during training.
ProAR is built upon Wan2.2-TI2V-5B.
- Code: github.com/luka-group/ProAR
- Evaluation data (for VBVR-10tasks): LinghuiShen/ProAR-test
- Paper: Coming soon
- Project page: Coming soon
Model Zoo
| Dataset | Model | Checkpoint | Training step |
|---|---|---|---|
| VBVR-10Tasks | AR | vbvr-10tasks/ar/model.pt |
10,000 |
| VBVR-10Tasks | ProAR | vbvr-10tasks/proar/model.pt |
10,000 |
| VideoRLVR-3Tasks | AR | videorlvr-3tasks/ar/model.pt |
5,000 |
| VideoRLVR-3Tasks | ProAR | videorlvr-3tasks/proar/model.pt |
5,000 |
| WorldArena | AR | worldarena/ar/model.pt |
10,000 |
| WorldArena | ProAR | worldarena/proar/model.pt |
10,000 |
Download
Download all released checkpoints:
hf download LinghuiShen/ProAR \
--local-dir checkpoints/ProAR
Model tree for LinghuiShen/ProAR
Base model
Wan-AI/Wan2.2-TI2V-5B