KBlueLeaf commited on
Commit
c9eef3f
·
verified ·
1 Parent(s): 24879fd

Model card: metadata, figures, links

Browse files
Files changed (5) hide show
  1. .gitattributes +2 -0
  2. README.md +153 -47
  3. assets/architecture.png +3 -0
  4. assets/overview.png +3 -0
  5. config.json +1 -2
.gitattributes CHANGED
@@ -33,3 +33,5 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ assets/architecture.png filter=lfs diff=lfs merge=lfs -text
37
+ assets/overview.png filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -4,36 +4,135 @@ library_name: transformers
4
  pipeline_tag: video-classification
5
  tags:
6
  - video
7
- - self-supervised
 
8
  - motion
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9
  ---
10
 
11
- # TT-VidT (TT3D, structured resample)
12
 
13
- Video encoder from *TT-VidT: Decoupling the Temporal Axis for Efficient
14
- Motion-Centric Video Pretraining*. A DINOv3 ViT-B/16 appearance path is paired with
15
- a compact Temporal Transfer (TT3D) pathway that produces per-frame motion tokens.
16
- TT3D down/up-samples the spatial tokens with a fixed structured weight plus a
17
- trainable channel mix: 195.5M parameters instead of 407.8M for the full-linear
18
- resample of the paper checkpoint.
19
 
20
- Code, training recipes and evaluation: see the TT-VidT repository.
21
 
22
- ## Usage
23
 
24
- With `transformers` only:
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
25
 
26
  ```python
27
  import torch
28
  from transformers import AutoModel
29
 
30
  model = AutoModel.from_pretrained("KBlueLeaf/TTVidT", trust_remote_code=True).cuda().eval()
31
- video = torch.rand(1, 8, 3, 256, 256, device="cuda") * 2 - 1 # [B, T, C, H, W], pixels in [-1, 1]
 
32
  with torch.no_grad(), torch.autocast("cuda", dtype=torch.float16):
33
- motion = model(video).motion_output # [B, T, 1, 768]
34
  ```
35
 
36
- With the TT-VidT repository:
 
37
 
38
  ```python
39
  from ttvidt.hub import load_model
@@ -42,35 +141,20 @@ model = load_model("KBlueLeaf/TTVidT", device="cuda")
42
  motion = model.encoder(video).motion_output
43
  ```
44
 
45
- Input: 8 frames at 256x256, normalised with mean = std = 0.5 (frames were sampled at
46
- 6 fps during training).
47
-
48
- ## Files
49
-
50
- | File | Content |
51
- |---|---|
52
- | `config.json` | architecture + `TTVidTrainer` arguments, and the `transformers` fields |
53
- | `model.safetensors` | fp32 encoder (195.5M parameters, incl. the 85.1M DINOv3 spatial path); no decoder |
54
- | `*.py` | encoder code for `trust_remote_code` (needs `torch` and `transformers` only) |
55
-
56
- Loading needs no access to the DINOv3 base weights.
57
-
58
- ## Training
59
-
60
- Initialised from the paper's TT3D + Diff Compression encoder (full-linear resample)
61
- with the resample replaced, then trained for 30k steps to match that encoder's motion
62
- and conditioning embeddings (relative MSE) on the pretraining data
63
- (OpenVid-1M 384 px + Moments-in-Time v2), batch 16, AdamW 5e-4 with muP
64
- scaling, 200 warmup steps, cosine decay.
65
 
66
  ## Results
67
 
68
- Frozen attentive probe, mean top-1 over 3 seeds (TT-VidT evaluation protocol, 8 frames):
 
 
 
 
 
69
 
70
- | Encoder | HMDB51 | ARID | IARD | Jester | SSv2 |
71
- |---|---:|---:|---:|---:|---:|
72
- | paper TT3D (full-linear resample, 407.8M) | 24.8 | 35.4 | 86.6 | 72.9 | 25.9 |
73
- | this checkpoint (structured resample, 195.5M) | 25.2 | 36.1 | 86.0 | 72.9 | 25.4 |
74
 
75
  ## Possible downstream uses
76
 
@@ -78,16 +162,38 @@ The encoder gives a compact sequence of per-frame motion tokens alongside the
78
  DINOv3 appearance features. Some directions we think are worth trying:
79
 
80
  - **From image models to video models**: pair an existing image model with the
81
- motion token sequence to get a video model for understanding in the broad
82
- sense: classification, retrieval, captioning, question answering, or any
83
- other task.
84
  - **Generation and motion transfer**: use the motion tokens as a conditioning
85
- signal for video generation, or take them from one clip and apply them to
86
- another subject or scene.
 
 
 
 
 
 
87
 
88
- Further exploration and feedback are very welcome, and so are attempts at larger
89
- scale (bigger backbones, more data, longer training).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
90
 
91
  ## License
92
 
93
- Apache-2.0. The spatial path is initialised from DINOv3, which has its own license.
 
 
 
4
  pipeline_tag: video-classification
5
  tags:
6
  - video
7
+ - video-representation-learning
8
+ - self-supervised-learning
9
  - motion
10
+ - temporal-modeling
11
+ - dinov3
12
+ - vision-transformer
13
+ - custom_code
14
+ base_model: facebook/dinov3-vitb16-pretrain-lvd1689m
15
+ base_model_relation: finetune
16
+ datasets:
17
+ - nkp37/OpenVid-1M
18
+ metrics:
19
+ - accuracy
20
+ model-index:
21
+ - name: TT-VidT (TT3D)
22
+ results:
23
+ - task:
24
+ type: video-classification
25
+ name: Frozen attentive probe
26
+ dataset:
27
+ type: hmdb51
28
+ name: HMDB51
29
+ metrics:
30
+ - type: accuracy
31
+ value: 25.2
32
+ name: Top-1 accuracy (mean of 3 seeds)
33
+ - task:
34
+ type: video-classification
35
+ name: Frozen attentive probe
36
+ dataset:
37
+ type: arid
38
+ name: ARID
39
+ metrics:
40
+ - type: accuracy
41
+ value: 36.1
42
+ name: Top-1 accuracy (mean of 3 seeds)
43
+ - task:
44
+ type: video-classification
45
+ name: Frozen attentive probe
46
+ dataset:
47
+ type: iard
48
+ name: IARD
49
+ metrics:
50
+ - type: accuracy
51
+ value: 86.0
52
+ name: Top-1 accuracy (mean of 3 seeds)
53
+ - task:
54
+ type: video-classification
55
+ name: Frozen attentive probe
56
+ dataset:
57
+ type: jester
58
+ name: Jester
59
+ metrics:
60
+ - type: accuracy
61
+ value: 72.9
62
+ name: Top-1 accuracy (mean of 3 seeds)
63
+ - task:
64
+ type: video-classification
65
+ name: Frozen attentive probe
66
+ dataset:
67
+ type: something-something-v2
68
+ name: Something-Something v2
69
+ metrics:
70
+ - type: accuracy
71
+ value: 25.4
72
+ name: Top-1 accuracy (mean of 3 seeds)
73
  ---
74
 
75
+ <div align="center">
76
 
77
+ # TT-VidT
 
 
 
 
 
78
 
79
+ ### Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
80
 
81
+ **NeurIPS 2026 (main track)**
82
 
83
+ [![Project Page](https://img.shields.io/badge/Project-Page-green)](https://kohakublueleaf.github.io/TTVidT/)
84
+ [![Paper](https://img.shields.io/badge/🤗%20Paper-2609.33419-yellow)](https://huggingface.co/papers/2609.33419)
85
+ [![arXiv](https://img.shields.io/badge/arXiv-2609.33419-b31b1b)](https://arxiv.org/abs/2609.33419)
86
+ [![Code](https://img.shields.io/badge/GitHub-KohakuBlueleaf%2FTTVidT-181717?logo=github)](https://github.com/KohakuBlueleaf/TTVidT)
87
+ [![Decoders](https://img.shields.io/badge/🤗%20Decoders-TTVidT--decoders-orange)](https://huggingface.co/KBlueLeaf/TTVidT-decoders)
88
+ [![License](https://img.shields.io/badge/License-Apache%202.0-blue)](#license)
89
+
90
+ </div>
91
+
92
+ ![TT-VidT overview](assets/overview.png)
93
+
94
+ **TT-VidT** is a self-supervised video encoder built for *motion*. A DINOv3 ViT-B/16
95
+ processes every frame independently (the appearance path), while a compact
96
+ **Temporal Transfer** pathway turns each frame into a few motion tokens that
97
+ exchange information across time. This repository holds the pretrained **TT3D**
98
+ encoder (195.5M parameters).
99
+
100
+ ## Model
101
+
102
+ ![TT-VidT architecture](assets/architecture.png)
103
+
104
+ - **Encoder**: 12 DINOv3 ViT-B/16 layers interleaved with 12 Temporal Transfer (TT3D)
105
+ layers. Each TT3D layer runs block-causal attention over the frame's K = 8 motion
106
+ tokens together with its 4x-downsampled spatial tokens, and writes the result back
107
+ to the spatial stream.
108
+ - **Pretraining objective**: Diff Compression. A DiT decoder reconstructs every
109
+ later frame from the *first* frame's features plus that frame's motion tokens, so
110
+ the motion tokens carry what the first frame cannot explain.
111
+
112
+ | | |
113
+ |---|---|
114
+ | Parameters | 195.5M (incl. the 85.1M DINOv3 spatial path) |
115
+ | Input | 8 RGB frames, 256 x 256, pixels in [-1, 1] (mean = std = 0.5), sampled at 6 fps in pretraining |
116
+ | Output | `motion_output`: one 768-d motion embedding per frame, `[B, T, 1, 768]` |
117
+ | Weights | fp32 `safetensors`, encoder only |
118
+
119
+ ## Quick start
120
+
121
+ **With `transformers` only** (the model code ships in this repository):
122
 
123
  ```python
124
  import torch
125
  from transformers import AutoModel
126
 
127
  model = AutoModel.from_pretrained("KBlueLeaf/TTVidT", trust_remote_code=True).cuda().eval()
128
+
129
+ video = torch.rand(1, 8, 3, 256, 256, device="cuda") * 2 - 1 # [B, T, C, H, W] in [-1, 1]
130
  with torch.no_grad(), torch.autocast("cuda", dtype=torch.float16):
131
+ motion = model(video).motion_output # [B, T, 1, 768]
132
  ```
133
 
134
+ **With the [TT-VidT codebase](https://github.com/KohakuBlueleaf/TTVidT)** (training,
135
+ evaluation, feature extraction):
136
 
137
  ```python
138
  from ttvidt.hub import load_model
 
141
  motion = model.encoder(video).motion_output
142
  ```
143
 
144
+ Loading needs no access to the (gated) DINOv3 base weights: every weight is in this
145
+ repository.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
146
 
147
  ## Results
148
 
149
+ Frozen attentive probe on the motion embeddings, top-1 accuracy (%), mean of 3 seeds,
150
+ 8 frames (evaluation protocol of the paper):
151
+
152
+ | HMDB51 | ARID | IARD | Jester | SSv2 |
153
+ |:---:|:---:|:---:|:---:|:---:|
154
+ | 25.2 | 36.1 | 86.0 | 72.9 | 25.4 |
155
 
156
+ For the full study (24 architecture–objective pairs at matched scale, fine-tuning and
157
+ diagnostics), see the [paper](https://huggingface.co/papers/2609.33419).
 
 
158
 
159
  ## Possible downstream uses
160
 
 
162
  DINOv3 appearance features. Some directions we think are worth trying:
163
 
164
  - **From image models to video models**: pair an existing image model with the
165
+ motion token sequence to get a video model for understanding in the broad sense:
166
+ classification, retrieval, captioning, question answering, or any other task.
 
167
  - **Generation and motion transfer**: use the motion tokens as a conditioning
168
+ signal for video generation, or take them from one clip and apply them to another
169
+ subject or scene.
170
+
171
+ > [!TIP]
172
+ > Further exploration and feedback are very welcome, and so are attempts at larger
173
+ > scale (bigger backbones, more data, longer training). Please open an issue or a
174
+ > discussion on [GitHub](https://github.com/KohakuBlueleaf/TTVidT) or in the
175
+ > Community tab.
176
 
177
+ ## Related resources
178
+
179
+ | Resource | Link |
180
+ |---|---|
181
+ | Project page | [kohakublueleaf.github.io/TTVidT](https://kohakublueleaf.github.io/TTVidT/) |
182
+ | Paper | [huggingface.co/papers/2609.33419](https://huggingface.co/papers/2609.33419) · [arXiv:2609.33419](https://arxiv.org/abs/2609.33419) |
183
+ | Source code (training, evaluation, all paper configs) | [github.com/KohakuBlueleaf/TTVidT](https://github.com/KohakuBlueleaf/TTVidT) |
184
+ | Pretrained DiT decoders (Diff Compression and the other objectives) | [KBlueLeaf/TTVidT-decoders](https://huggingface.co/KBlueLeaf/TTVidT-decoders) |
185
+
186
+ ## Files
187
+
188
+ | File | Content |
189
+ |---|---|
190
+ | `config.json` | architecture and loader configuration (`transformers` + TT-VidT codebase) |
191
+ | `model.safetensors` | fp32 encoder weights |
192
+ | `*.py` | encoder code for `trust_remote_code` (needs only `torch` and `transformers`) |
193
+ | `assets/` | figures of this card |
194
 
195
  ## License
196
 
197
+ Apache-2.0. The spatial path is initialised from
198
+ [DINOv3](https://huggingface.co/facebook/dinov3-vitb16-pretrain-lvd1689m), which is
199
+ released under its own license.
assets/architecture.png ADDED

Git LFS Details

  • SHA256: dc88bbbe3ccf23f4bed9a6b2a1121985c3bd9dd507070cbe22feaa7e90722b79
  • Pointer size: 131 Bytes
  • Size of remote file: 280 kB
assets/overview.png ADDED

Git LFS Details

  • SHA256: 2ccb9be3afc8a270595d10746a9d40bc7801f0172ab8c348e2d2595884bca2fe
  • Pointer size: 132 Bytes
  • Size of remote file: 2.52 MB
config.json CHANGED
@@ -158,8 +158,7 @@
158
  }
159
  },
160
  "encoder_only": true,
161
- "name": "ttvidt-tt3d-structured",
162
- "source": "distilled from ttvidt-tt3d-diffcomp",
163
  "model_type": "ttvidt",
164
  "architectures": [
165
  "TTVidTModel"
 
158
  }
159
  },
160
  "encoder_only": true,
161
+ "name": "ttvidt-tt3d",
 
162
  "model_type": "ttvidt",
163
  "architectures": [
164
  "TTVidTModel"