Update README.md
Browse files
README.md
CHANGED
|
@@ -6,7 +6,6 @@ tags:
|
|
| 6 |
- inverse-dynamics
|
| 7 |
- camera-motion
|
| 8 |
- optical-flow
|
| 9 |
-
- counter-strike-2
|
| 10 |
model-index:
|
| 11 |
- name: RIDM flow model
|
| 12 |
results:
|
|
@@ -33,9 +32,6 @@ model-index:
|
|
| 33 |
- task:
|
| 34 |
type: video-classification
|
| 35 |
name: Forward, turn left or turn right
|
| 36 |
-
dataset:
|
| 37 |
-
name: Counter-Strike 2, 7,716 frame pairs
|
| 38 |
-
type: counter-strike-2
|
| 39 |
metrics:
|
| 40 |
- type: accuracy
|
| 41 |
value: 84.5
|
|
@@ -43,9 +39,6 @@ model-index:
|
|
| 43 |
- task:
|
| 44 |
type: video-classification
|
| 45 |
name: Turn direction
|
| 46 |
-
dataset:
|
| 47 |
-
name: Counter-Strike 2, 3,716 turn pairs
|
| 48 |
-
type: counter-strike-2
|
| 49 |
metrics:
|
| 50 |
- type: accuracy
|
| 51 |
value: 84.9
|
|
@@ -55,18 +48,18 @@ model-index:
|
|
| 55 |
# Reka Inverse Dynamics Model (RIDM)
|
| 56 |
|
| 57 |
Two models predict camera motion (W/A/S/D/Shift, yaw, pitch) from video.
|
| 58 |
-
Both trained on
|
| 59 |
|
| 60 |
| | [flow](flow/) | [pixel](pixel/) |
|
| 61 |
|---|---|---|
|
| 62 |
| Input | Optical flow (RAFT-small) | Raw frames |
|
| 63 |
| Parameters | 1,795,337 trained, plus 990,162 in frozen RAFT-small (2,785,499 in total) | 9,836,063 |
|
| 64 |
| Output | One prediction per 17-frame window | One prediction per frame |
|
| 65 |
-
| Real-video turn accuracy | 91.5 | 51.1
|
| 66 |
-
|
|
| 67 |
| Dependencies | torch, torchvision, av, opencv, safetensors | torch, av, opencv, safetensors |
|
| 68 |
|
| 69 |
-
|
| 70 |
|
| 71 |
```text
|
| 72 |
Reka-Inverse-Dynamics-Model/
|
|
@@ -88,8 +81,8 @@ In the failure clip the camera stands still and pans. The model reports the walk
|
|
| 88 |
|
| 89 |
## Which model to use
|
| 90 |
|
| 91 |
-
|
| 92 |
-
|
| 93 |
|
| 94 |
## Use
|
| 95 |
|
|
@@ -110,7 +103,7 @@ Status: preliminary release. The scores marked † are unconfirmed.
|
|
| 110 |
|
| 111 |
## Training data
|
| 112 |
|
| 113 |
-
- Source: 4,071 recorded
|
| 114 |
- Labels: the game engine gives the keys, the mouse movement and the camera angles for every frame. Nobody annotates.
|
| 115 |
- Flow model: 7,071,037 frame pairs. Pixel model: about 104,000 gameplay windows.
|
| 116 |
- Both models trained on one NVIDIA L4 GPU.
|
|
@@ -119,7 +112,7 @@ Status: preliminary release. The scores marked † are unconfirmed.
|
|
| 119 |
|
| 120 |
- The weights and the code in this repository use the Apache License 2.0. See `LICENSE`.
|
| 121 |
- RAFT-small weights use the BSD-3 license. They come from torchvision and are not in this repository. See `NOTICE`.
|
| 122 |
-
- The models trained on
|
| 123 |
- The videos in `assets/` and `examples/` are not Apache 2.0. They come from walking tours by POPtravel on Wikimedia Commons, licensed CC BY 3.0, and carry the model overlay or are shortened excerpts.
|
| 124 |
Sources: [Bristol](https://commons.wikimedia.org/wiki/File:Walking_in_BRISTOL_-_UK_-_4K_60fps_(UHD).webm),
|
| 125 |
[Ingolstadt](https://commons.wikimedia.org/wiki/File:Walking_in_INGOLSTADT_-_Germany_-_4K_60fps_(UHD).webm),
|
|
|
|
| 6 |
- inverse-dynamics
|
| 7 |
- camera-motion
|
| 8 |
- optical-flow
|
|
|
|
| 9 |
model-index:
|
| 10 |
- name: RIDM flow model
|
| 11 |
results:
|
|
|
|
| 32 |
- task:
|
| 33 |
type: video-classification
|
| 34 |
name: Forward, turn left or turn right
|
|
|
|
|
|
|
|
|
|
| 35 |
metrics:
|
| 36 |
- type: accuracy
|
| 37 |
value: 84.5
|
|
|
|
| 39 |
- task:
|
| 40 |
type: video-classification
|
| 41 |
name: Turn direction
|
|
|
|
|
|
|
|
|
|
| 42 |
metrics:
|
| 43 |
- type: accuracy
|
| 44 |
value: 84.9
|
|
|
|
| 48 |
# Reka Inverse Dynamics Model (RIDM)
|
| 49 |
|
| 50 |
Two models predict camera motion (W/A/S/D/Shift, yaw, pitch) from video.
|
| 51 |
+
Both trained on game renders.
|
| 52 |
|
| 53 |
| | [flow](flow/) | [pixel](pixel/) |
|
| 54 |
|---|---|---|
|
| 55 |
| Input | Optical flow (RAFT-small) | Raw frames |
|
| 56 |
| Parameters | 1,795,337 trained, plus 990,162 in frozen RAFT-small (2,785,499 in total) | 9,836,063 |
|
| 57 |
| Output | One prediction per 17-frame window | One prediction per frame |
|
| 58 |
+
| Real-video turn accuracy | 91.5 | 51.1 |
|
| 59 |
+
| Gaming footage turn accuracy | 84.9 | 70.9 |
|
| 60 |
| Dependencies | torch, torchvision, av, opencv, safetensors | torch, av, opencv, safetensors |
|
| 61 |
|
| 62 |
+
All scores are in `results.json`.
|
| 63 |
|
| 64 |
```text
|
| 65 |
Reka-Inverse-Dynamics-Model/
|
|
|
|
| 81 |
|
| 82 |
## Which model to use
|
| 83 |
|
| 84 |
+
The flow model has better generalization capabilities to real videos.
|
| 85 |
+
The pixel model might have a high potential for gaming footage especially if trained further.
|
| 86 |
|
| 87 |
## Use
|
| 88 |
|
|
|
|
| 103 |
|
| 104 |
## Training data
|
| 105 |
|
| 106 |
+
- Source: 4,071 recorded gaming footage plays, captured at 1,280 x 720 and 48 frames per second.
|
| 107 |
- Labels: the game engine gives the keys, the mouse movement and the camera angles for every frame. Nobody annotates.
|
| 108 |
- Flow model: 7,071,037 frame pairs. Pixel model: about 104,000 gameplay windows.
|
| 109 |
- Both models trained on one NVIDIA L4 GPU.
|
|
|
|
| 112 |
|
| 113 |
- The weights and the code in this repository use the Apache License 2.0. See `LICENSE`.
|
| 114 |
- RAFT-small weights use the BSD-3 license. They come from torchvision and are not in this repository. See `NOTICE`.
|
| 115 |
+
- The models trained on gameplay captures. This repository holds no game assets, no gameplay clips and no training data.
|
| 116 |
- The videos in `assets/` and `examples/` are not Apache 2.0. They come from walking tours by POPtravel on Wikimedia Commons, licensed CC BY 3.0, and carry the model overlay or are shortened excerpts.
|
| 117 |
Sources: [Bristol](https://commons.wikimedia.org/wiki/File:Walking_in_BRISTOL_-_UK_-_4K_60fps_(UHD).webm),
|
| 118 |
[Ingolstadt](https://commons.wikimedia.org/wiki/File:Walking_in_INGOLSTADT_-_Germany_-_4K_60fps_(UHD).webm),
|