fedorzh commited on
Commit
94070e8
·
verified ·
1 Parent(s): 97a0336

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +8 -15
README.md CHANGED
@@ -6,7 +6,6 @@ tags:
6
  - inverse-dynamics
7
  - camera-motion
8
  - optical-flow
9
- - counter-strike-2
10
  model-index:
11
  - name: RIDM flow model
12
  results:
@@ -33,9 +32,6 @@ model-index:
33
  - task:
34
  type: video-classification
35
  name: Forward, turn left or turn right
36
- dataset:
37
- name: Counter-Strike 2, 7,716 frame pairs
38
- type: counter-strike-2
39
  metrics:
40
  - type: accuracy
41
  value: 84.5
@@ -43,9 +39,6 @@ model-index:
43
  - task:
44
  type: video-classification
45
  name: Turn direction
46
- dataset:
47
- name: Counter-Strike 2, 3,716 turn pairs
48
- type: counter-strike-2
49
  metrics:
50
  - type: accuracy
51
  value: 84.9
@@ -55,18 +48,18 @@ model-index:
55
  # Reka Inverse Dynamics Model (RIDM)
56
 
57
  Two models predict camera motion (W/A/S/D/Shift, yaw, pitch) from video.
58
- Both trained on Counter-Strike 2 renders.
59
 
60
  | | [flow](flow/) | [pixel](pixel/) |
61
  |---|---|---|
62
  | Input | Optical flow (RAFT-small) | Raw frames |
63
  | Parameters | 1,795,337 trained, plus 990,162 in frozen RAFT-small (2,785,499 in total) | 9,836,063 |
64
  | Output | One prediction per 17-frame window | One prediction per frame |
65
- | Real-video turn accuracy | 91.5 | 51.1 † |
66
- | Counter-Strike 2 turn accuracy | 84.9 | 70.9 † |
67
  | Dependencies | torch, torchvision, av, opencv, safetensors | torch, av, opencv, safetensors |
68
 
69
- † The number comes from a write-up and has no run log yet. All scores are in `results.json`.
70
 
71
  ```text
72
  Reka-Inverse-Dynamics-Model/
@@ -88,8 +81,8 @@ In the failure clip the camera stands still and pans. The model reports the walk
88
 
89
  ## Which model to use
90
 
91
- Use the flow model for real video. The pixel model gives no clear signal on real footage.
92
- Its card says so. Use the pixel model for Counter-Strike 2 footage only.
93
 
94
  ## Use
95
 
@@ -110,7 +103,7 @@ Status: preliminary release. The scores marked † are unconfirmed.
110
 
111
  ## Training data
112
 
113
- - Source: 4,071 recorded Counter-Strike 2 matches, captured at 1,280 x 720 and 48 frames per second.
114
  - Labels: the game engine gives the keys, the mouse movement and the camera angles for every frame. Nobody annotates.
115
  - Flow model: 7,071,037 frame pairs. Pixel model: about 104,000 gameplay windows.
116
  - Both models trained on one NVIDIA L4 GPU.
@@ -119,7 +112,7 @@ Status: preliminary release. The scores marked † are unconfirmed.
119
 
120
  - The weights and the code in this repository use the Apache License 2.0. See `LICENSE`.
121
  - RAFT-small weights use the BSD-3 license. They come from torchvision and are not in this repository. See `NOTICE`.
122
- - The models trained on Counter-Strike 2 gameplay captures. This repository holds no game assets, no gameplay clips and no training data.
123
  - The videos in `assets/` and `examples/` are not Apache 2.0. They come from walking tours by POPtravel on Wikimedia Commons, licensed CC BY 3.0, and carry the model overlay or are shortened excerpts.
124
  Sources: [Bristol](https://commons.wikimedia.org/wiki/File:Walking_in_BRISTOL_-_UK_-_4K_60fps_(UHD).webm),
125
  [Ingolstadt](https://commons.wikimedia.org/wiki/File:Walking_in_INGOLSTADT_-_Germany_-_4K_60fps_(UHD).webm),
 
6
  - inverse-dynamics
7
  - camera-motion
8
  - optical-flow
 
9
  model-index:
10
  - name: RIDM flow model
11
  results:
 
32
  - task:
33
  type: video-classification
34
  name: Forward, turn left or turn right
 
 
 
35
  metrics:
36
  - type: accuracy
37
  value: 84.5
 
39
  - task:
40
  type: video-classification
41
  name: Turn direction
 
 
 
42
  metrics:
43
  - type: accuracy
44
  value: 84.9
 
48
  # Reka Inverse Dynamics Model (RIDM)
49
 
50
  Two models predict camera motion (W/A/S/D/Shift, yaw, pitch) from video.
51
+ Both trained on game renders.
52
 
53
  | | [flow](flow/) | [pixel](pixel/) |
54
  |---|---|---|
55
  | Input | Optical flow (RAFT-small) | Raw frames |
56
  | Parameters | 1,795,337 trained, plus 990,162 in frozen RAFT-small (2,785,499 in total) | 9,836,063 |
57
  | Output | One prediction per 17-frame window | One prediction per frame |
58
+ | Real-video turn accuracy | 91.5 | 51.1 |
59
+ | Gaming footage turn accuracy | 84.9 | 70.9 |
60
  | Dependencies | torch, torchvision, av, opencv, safetensors | torch, av, opencv, safetensors |
61
 
62
+ All scores are in `results.json`.
63
 
64
  ```text
65
  Reka-Inverse-Dynamics-Model/
 
81
 
82
  ## Which model to use
83
 
84
+ The flow model has better generalization capabilities to real videos.
85
+ The pixel model might have a high potential for gaming footage especially if trained further.
86
 
87
  ## Use
88
 
 
103
 
104
  ## Training data
105
 
106
+ - Source: 4,071 recorded gaming footage plays, captured at 1,280 x 720 and 48 frames per second.
107
  - Labels: the game engine gives the keys, the mouse movement and the camera angles for every frame. Nobody annotates.
108
  - Flow model: 7,071,037 frame pairs. Pixel model: about 104,000 gameplay windows.
109
  - Both models trained on one NVIDIA L4 GPU.
 
112
 
113
  - The weights and the code in this repository use the Apache License 2.0. See `LICENSE`.
114
  - RAFT-small weights use the BSD-3 license. They come from torchvision and are not in this repository. See `NOTICE`.
115
+ - The models trained on gameplay captures. This repository holds no game assets, no gameplay clips and no training data.
116
  - The videos in `assets/` and `examples/` are not Apache 2.0. They come from walking tours by POPtravel on Wikimedia Commons, licensed CC BY 3.0, and carry the model overlay or are shortened excerpts.
117
  Sources: [Bristol](https://commons.wikimedia.org/wiki/File:Walking_in_BRISTOL_-_UK_-_4K_60fps_(UHD).webm),
118
  [Ingolstadt](https://commons.wikimedia.org/wiki/File:Walking_in_INGOLSTADT_-_Germany_-_4K_60fps_(UHD).webm),