Video Classification
Transformers
Safetensors
ttvidt
feature-extraction
video
video-representation-learning
self-supervised-learning
motion
temporal-modeling
dinov3
vision-transformer
custom_code
Eval Results (legacy)
Instructions to use KBlueLeaf/TTVidT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use KBlueLeaf/TTVidT with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("video-classification", model="KBlueLeaf/TTVidT", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("KBlueLeaf/TTVidT", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Model card: metadata, figures, links
Browse files- .gitattributes +2 -0
- README.md +153 -47
- assets/architecture.png +3 -0
- assets/overview.png +3 -0
- config.json +1 -2
.gitattributes
CHANGED
|
@@ -33,3 +33,5 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
assets/architecture.png filter=lfs diff=lfs merge=lfs -text
|
| 37 |
+
assets/overview.png filter=lfs diff=lfs merge=lfs -text
|
README.md
CHANGED
|
@@ -4,36 +4,135 @@ library_name: transformers
|
|
| 4 |
pipeline_tag: video-classification
|
| 5 |
tags:
|
| 6 |
- video
|
| 7 |
-
-
|
|
|
|
| 8 |
- motion
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 9 |
---
|
| 10 |
|
| 11 |
-
|
| 12 |
|
| 13 |
-
|
| 14 |
-
Motion-Centric Video Pretraining*. A DINOv3 ViT-B/16 appearance path is paired with
|
| 15 |
-
a compact Temporal Transfer (TT3D) pathway that produces per-frame motion tokens.
|
| 16 |
-
TT3D down/up-samples the spatial tokens with a fixed structured weight plus a
|
| 17 |
-
trainable channel mix: 195.5M parameters instead of 407.8M for the full-linear
|
| 18 |
-
resample of the paper checkpoint.
|
| 19 |
|
| 20 |
-
|
| 21 |
|
| 22 |
-
|
| 23 |
|
| 24 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 25 |
|
| 26 |
```python
|
| 27 |
import torch
|
| 28 |
from transformers import AutoModel
|
| 29 |
|
| 30 |
model = AutoModel.from_pretrained("KBlueLeaf/TTVidT", trust_remote_code=True).cuda().eval()
|
| 31 |
-
|
|
|
|
| 32 |
with torch.no_grad(), torch.autocast("cuda", dtype=torch.float16):
|
| 33 |
-
motion = model(video).motion_output
|
| 34 |
```
|
| 35 |
|
| 36 |
-
With the TT-VidT
|
|
|
|
| 37 |
|
| 38 |
```python
|
| 39 |
from ttvidt.hub import load_model
|
|
@@ -42,35 +141,20 @@ model = load_model("KBlueLeaf/TTVidT", device="cuda")
|
|
| 42 |
motion = model.encoder(video).motion_output
|
| 43 |
```
|
| 44 |
|
| 45 |
-
|
| 46 |
-
|
| 47 |
-
|
| 48 |
-
## Files
|
| 49 |
-
|
| 50 |
-
| File | Content |
|
| 51 |
-
|---|---|
|
| 52 |
-
| `config.json` | architecture + `TTVidTrainer` arguments, and the `transformers` fields |
|
| 53 |
-
| `model.safetensors` | fp32 encoder (195.5M parameters, incl. the 85.1M DINOv3 spatial path); no decoder |
|
| 54 |
-
| `*.py` | encoder code for `trust_remote_code` (needs `torch` and `transformers` only) |
|
| 55 |
-
|
| 56 |
-
Loading needs no access to the DINOv3 base weights.
|
| 57 |
-
|
| 58 |
-
## Training
|
| 59 |
-
|
| 60 |
-
Initialised from the paper's TT3D + Diff Compression encoder (full-linear resample)
|
| 61 |
-
with the resample replaced, then trained for 30k steps to match that encoder's motion
|
| 62 |
-
and conditioning embeddings (relative MSE) on the pretraining data
|
| 63 |
-
(OpenVid-1M 384 px + Moments-in-Time v2), batch 16, AdamW 5e-4 with muP
|
| 64 |
-
scaling, 200 warmup steps, cosine decay.
|
| 65 |
|
| 66 |
## Results
|
| 67 |
|
| 68 |
-
Frozen attentive probe
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 69 |
|
| 70 |
-
|
| 71 |
-
|
| 72 |
-
| paper TT3D (full-linear resample, 407.8M) | 24.8 | 35.4 | 86.6 | 72.9 | 25.9 |
|
| 73 |
-
| this checkpoint (structured resample, 195.5M) | 25.2 | 36.1 | 86.0 | 72.9 | 25.4 |
|
| 74 |
|
| 75 |
## Possible downstream uses
|
| 76 |
|
|
@@ -78,16 +162,38 @@ The encoder gives a compact sequence of per-frame motion tokens alongside the
|
|
| 78 |
DINOv3 appearance features. Some directions we think are worth trying:
|
| 79 |
|
| 80 |
- **From image models to video models**: pair an existing image model with the
|
| 81 |
-
motion token sequence to get a video model for understanding in the broad
|
| 82 |
-
|
| 83 |
-
other task.
|
| 84 |
- **Generation and motion transfer**: use the motion tokens as a conditioning
|
| 85 |
-
signal for video generation, or take them from one clip and apply them to
|
| 86 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 87 |
|
| 88 |
-
|
| 89 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 90 |
|
| 91 |
## License
|
| 92 |
|
| 93 |
-
Apache-2.0. The spatial path is initialised from
|
|
|
|
|
|
|
|
|
| 4 |
pipeline_tag: video-classification
|
| 5 |
tags:
|
| 6 |
- video
|
| 7 |
+
- video-representation-learning
|
| 8 |
+
- self-supervised-learning
|
| 9 |
- motion
|
| 10 |
+
- temporal-modeling
|
| 11 |
+
- dinov3
|
| 12 |
+
- vision-transformer
|
| 13 |
+
- custom_code
|
| 14 |
+
base_model: facebook/dinov3-vitb16-pretrain-lvd1689m
|
| 15 |
+
base_model_relation: finetune
|
| 16 |
+
datasets:
|
| 17 |
+
- nkp37/OpenVid-1M
|
| 18 |
+
metrics:
|
| 19 |
+
- accuracy
|
| 20 |
+
model-index:
|
| 21 |
+
- name: TT-VidT (TT3D)
|
| 22 |
+
results:
|
| 23 |
+
- task:
|
| 24 |
+
type: video-classification
|
| 25 |
+
name: Frozen attentive probe
|
| 26 |
+
dataset:
|
| 27 |
+
type: hmdb51
|
| 28 |
+
name: HMDB51
|
| 29 |
+
metrics:
|
| 30 |
+
- type: accuracy
|
| 31 |
+
value: 25.2
|
| 32 |
+
name: Top-1 accuracy (mean of 3 seeds)
|
| 33 |
+
- task:
|
| 34 |
+
type: video-classification
|
| 35 |
+
name: Frozen attentive probe
|
| 36 |
+
dataset:
|
| 37 |
+
type: arid
|
| 38 |
+
name: ARID
|
| 39 |
+
metrics:
|
| 40 |
+
- type: accuracy
|
| 41 |
+
value: 36.1
|
| 42 |
+
name: Top-1 accuracy (mean of 3 seeds)
|
| 43 |
+
- task:
|
| 44 |
+
type: video-classification
|
| 45 |
+
name: Frozen attentive probe
|
| 46 |
+
dataset:
|
| 47 |
+
type: iard
|
| 48 |
+
name: IARD
|
| 49 |
+
metrics:
|
| 50 |
+
- type: accuracy
|
| 51 |
+
value: 86.0
|
| 52 |
+
name: Top-1 accuracy (mean of 3 seeds)
|
| 53 |
+
- task:
|
| 54 |
+
type: video-classification
|
| 55 |
+
name: Frozen attentive probe
|
| 56 |
+
dataset:
|
| 57 |
+
type: jester
|
| 58 |
+
name: Jester
|
| 59 |
+
metrics:
|
| 60 |
+
- type: accuracy
|
| 61 |
+
value: 72.9
|
| 62 |
+
name: Top-1 accuracy (mean of 3 seeds)
|
| 63 |
+
- task:
|
| 64 |
+
type: video-classification
|
| 65 |
+
name: Frozen attentive probe
|
| 66 |
+
dataset:
|
| 67 |
+
type: something-something-v2
|
| 68 |
+
name: Something-Something v2
|
| 69 |
+
metrics:
|
| 70 |
+
- type: accuracy
|
| 71 |
+
value: 25.4
|
| 72 |
+
name: Top-1 accuracy (mean of 3 seeds)
|
| 73 |
---
|
| 74 |
|
| 75 |
+
<div align="center">
|
| 76 |
|
| 77 |
+
# TT-VidT
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 78 |
|
| 79 |
+
### Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
|
| 80 |
|
| 81 |
+
**NeurIPS 2026 (main track)**
|
| 82 |
|
| 83 |
+
[](https://kohakublueleaf.github.io/TTVidT/)
|
| 84 |
+
[](https://huggingface.co/papers/2609.33419)
|
| 85 |
+
[](https://arxiv.org/abs/2609.33419)
|
| 86 |
+
[](https://github.com/KohakuBlueleaf/TTVidT)
|
| 87 |
+
[](https://huggingface.co/KBlueLeaf/TTVidT-decoders)
|
| 88 |
+
[](#license)
|
| 89 |
+
|
| 90 |
+
</div>
|
| 91 |
+
|
| 92 |
+

|
| 93 |
+
|
| 94 |
+
**TT-VidT** is a self-supervised video encoder built for *motion*. A DINOv3 ViT-B/16
|
| 95 |
+
processes every frame independently (the appearance path), while a compact
|
| 96 |
+
**Temporal Transfer** pathway turns each frame into a few motion tokens that
|
| 97 |
+
exchange information across time. This repository holds the pretrained **TT3D**
|
| 98 |
+
encoder (195.5M parameters).
|
| 99 |
+
|
| 100 |
+
## Model
|
| 101 |
+
|
| 102 |
+

|
| 103 |
+
|
| 104 |
+
- **Encoder**: 12 DINOv3 ViT-B/16 layers interleaved with 12 Temporal Transfer (TT3D)
|
| 105 |
+
layers. Each TT3D layer runs block-causal attention over the frame's K = 8 motion
|
| 106 |
+
tokens together with its 4x-downsampled spatial tokens, and writes the result back
|
| 107 |
+
to the spatial stream.
|
| 108 |
+
- **Pretraining objective**: Diff Compression. A DiT decoder reconstructs every
|
| 109 |
+
later frame from the *first* frame's features plus that frame's motion tokens, so
|
| 110 |
+
the motion tokens carry what the first frame cannot explain.
|
| 111 |
+
|
| 112 |
+
| | |
|
| 113 |
+
|---|---|
|
| 114 |
+
| Parameters | 195.5M (incl. the 85.1M DINOv3 spatial path) |
|
| 115 |
+
| Input | 8 RGB frames, 256 x 256, pixels in [-1, 1] (mean = std = 0.5), sampled at 6 fps in pretraining |
|
| 116 |
+
| Output | `motion_output`: one 768-d motion embedding per frame, `[B, T, 1, 768]` |
|
| 117 |
+
| Weights | fp32 `safetensors`, encoder only |
|
| 118 |
+
|
| 119 |
+
## Quick start
|
| 120 |
+
|
| 121 |
+
**With `transformers` only** (the model code ships in this repository):
|
| 122 |
|
| 123 |
```python
|
| 124 |
import torch
|
| 125 |
from transformers import AutoModel
|
| 126 |
|
| 127 |
model = AutoModel.from_pretrained("KBlueLeaf/TTVidT", trust_remote_code=True).cuda().eval()
|
| 128 |
+
|
| 129 |
+
video = torch.rand(1, 8, 3, 256, 256, device="cuda") * 2 - 1 # [B, T, C, H, W] in [-1, 1]
|
| 130 |
with torch.no_grad(), torch.autocast("cuda", dtype=torch.float16):
|
| 131 |
+
motion = model(video).motion_output # [B, T, 1, 768]
|
| 132 |
```
|
| 133 |
|
| 134 |
+
**With the [TT-VidT codebase](https://github.com/KohakuBlueleaf/TTVidT)** (training,
|
| 135 |
+
evaluation, feature extraction):
|
| 136 |
|
| 137 |
```python
|
| 138 |
from ttvidt.hub import load_model
|
|
|
|
| 141 |
motion = model.encoder(video).motion_output
|
| 142 |
```
|
| 143 |
|
| 144 |
+
Loading needs no access to the (gated) DINOv3 base weights: every weight is in this
|
| 145 |
+
repository.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 146 |
|
| 147 |
## Results
|
| 148 |
|
| 149 |
+
Frozen attentive probe on the motion embeddings, top-1 accuracy (%), mean of 3 seeds,
|
| 150 |
+
8 frames (evaluation protocol of the paper):
|
| 151 |
+
|
| 152 |
+
| HMDB51 | ARID | IARD | Jester | SSv2 |
|
| 153 |
+
|:---:|:---:|:---:|:---:|:---:|
|
| 154 |
+
| 25.2 | 36.1 | 86.0 | 72.9 | 25.4 |
|
| 155 |
|
| 156 |
+
For the full study (24 architecture–objective pairs at matched scale, fine-tuning and
|
| 157 |
+
diagnostics), see the [paper](https://huggingface.co/papers/2609.33419).
|
|
|
|
|
|
|
| 158 |
|
| 159 |
## Possible downstream uses
|
| 160 |
|
|
|
|
| 162 |
DINOv3 appearance features. Some directions we think are worth trying:
|
| 163 |
|
| 164 |
- **From image models to video models**: pair an existing image model with the
|
| 165 |
+
motion token sequence to get a video model for understanding in the broad sense:
|
| 166 |
+
classification, retrieval, captioning, question answering, or any other task.
|
|
|
|
| 167 |
- **Generation and motion transfer**: use the motion tokens as a conditioning
|
| 168 |
+
signal for video generation, or take them from one clip and apply them to another
|
| 169 |
+
subject or scene.
|
| 170 |
+
|
| 171 |
+
> [!TIP]
|
| 172 |
+
> Further exploration and feedback are very welcome, and so are attempts at larger
|
| 173 |
+
> scale (bigger backbones, more data, longer training). Please open an issue or a
|
| 174 |
+
> discussion on [GitHub](https://github.com/KohakuBlueleaf/TTVidT) or in the
|
| 175 |
+
> Community tab.
|
| 176 |
|
| 177 |
+
## Related resources
|
| 178 |
+
|
| 179 |
+
| Resource | Link |
|
| 180 |
+
|---|---|
|
| 181 |
+
| Project page | [kohakublueleaf.github.io/TTVidT](https://kohakublueleaf.github.io/TTVidT/) |
|
| 182 |
+
| Paper | [huggingface.co/papers/2609.33419](https://huggingface.co/papers/2609.33419) · [arXiv:2609.33419](https://arxiv.org/abs/2609.33419) |
|
| 183 |
+
| Source code (training, evaluation, all paper configs) | [github.com/KohakuBlueleaf/TTVidT](https://github.com/KohakuBlueleaf/TTVidT) |
|
| 184 |
+
| Pretrained DiT decoders (Diff Compression and the other objectives) | [KBlueLeaf/TTVidT-decoders](https://huggingface.co/KBlueLeaf/TTVidT-decoders) |
|
| 185 |
+
|
| 186 |
+
## Files
|
| 187 |
+
|
| 188 |
+
| File | Content |
|
| 189 |
+
|---|---|
|
| 190 |
+
| `config.json` | architecture and loader configuration (`transformers` + TT-VidT codebase) |
|
| 191 |
+
| `model.safetensors` | fp32 encoder weights |
|
| 192 |
+
| `*.py` | encoder code for `trust_remote_code` (needs only `torch` and `transformers`) |
|
| 193 |
+
| `assets/` | figures of this card |
|
| 194 |
|
| 195 |
## License
|
| 196 |
|
| 197 |
+
Apache-2.0. The spatial path is initialised from
|
| 198 |
+
[DINOv3](https://huggingface.co/facebook/dinov3-vitb16-pretrain-lvd1689m), which is
|
| 199 |
+
released under its own license.
|
assets/architecture.png
ADDED
|
Git LFS Details
|
assets/overview.png
ADDED
|
Git LFS Details
|
config.json
CHANGED
|
@@ -158,8 +158,7 @@
|
|
| 158 |
}
|
| 159 |
},
|
| 160 |
"encoder_only": true,
|
| 161 |
-
"name": "ttvidt-tt3d
|
| 162 |
-
"source": "distilled from ttvidt-tt3d-diffcomp",
|
| 163 |
"model_type": "ttvidt",
|
| 164 |
"architectures": [
|
| 165 |
"TTVidTModel"
|
|
|
|
| 158 |
}
|
| 159 |
},
|
| 160 |
"encoder_only": true,
|
| 161 |
+
"name": "ttvidt-tt3d",
|
|
|
|
| 162 |
"model_type": "ttvidt",
|
| 163 |
"architectures": [
|
| 164 |
"TTVidTModel"
|