File size: 10,906 Bytes
700dd75
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
# Training Guide

This guide covers data processing, training, evaluation, and ONNX export for
SONIC whole-body controllers.

## Overview

SONIC uses a universal-token architecture to control a humanoid robot (Unitree
G1, 29 DOF) by imitating human motion capture data. Multiple parallel encoders
accept different motion input formats:

- **G1**: Robot joint trajectories
- **Teleop**: VR 3-point tracking targets (head + two wrists)
- **SMPL**: Parametric human body model joint positions
- **SOMA**: BVH-derived skeleton joint positions (optional 4th encoder)

All encoders project into a shared latent token space via FSQ (Finite Scalar
Quantization), and a single decoder produces joint actions regardless of input
modality. Training uses PPO with auxiliary losses in Isaac Lab simulation.

| Config | Encoders | Use case |
|--------|----------|----------|
| `sonic_release` | G1, teleop, SMPL | **Default** β€” matches the released checkpoint |
| `sonic_v1_1` | G1, teleop, SMPL | SONIC v1.1 with heading-normalized targets and wrist-pose augmentation |
| `sonic_bones_seed` | G1, teleop, SMPL, SOMA | Extended training with SOMA skeleton encoder |

Use `sonic_release` for finetuning and evaluation. The `sonic_bones_seed`
config adds a fourth SOMA encoder (see [Training with SOMA](#training-with-soma-encoder)).
Use `sonic_v1_1` with `sonic_v1_1/last.pt`.

## Data Processing

### Step 1: Convert motion data

SONIC requires motion data in **motion_lib PKL format**. Convert Bones-SEED
CSV files:

```bash
python gear_sonic/data_process/convert_soma_csv_to_motion_lib.py \
    --input /path/to/bones_seed/g1/csv/ \
    --output data/motion_lib_bones_seed/robot \
    --fps 30 \
    --fps_source 120 \
    --individual \
    --num_workers 16
```

### Step 2: Filter motions

Remove motions the G1 robot cannot perform (furniture interaction, vehicles,
acrobatics, elevated surfaces):

```bash
python gear_sonic/data_process/filter_and_copy_bones_data.py \
    --source data/motion_lib_bones_seed/robot \
    --dest data/motion_lib_bones_seed/robot_filtered \
    --workers 16
```

This removes ~8.7% of motions (~130K of 142K remain). Use `--dry-run` to
preview, or `--add-keywords` to add custom filters.

### Data layout

Place processed data at the repo root:

```
<repo_root>/
β”œβ”€β”€ data/motion_lib_bones_seed/
β”‚   β”œβ”€β”€ robot/              # Full motion library (142K PKLs)
β”‚   └── robot_filtered/     # Filtered subset (~130K PKLs)
└── gear_sonic/
```

## Training

### Basic command

```bash
python gear_sonic/train_agent_trl.py \
    +exp=manager/universal_token/all_modes/sonic_release \
    num_envs=4096 headless=True \
    ++manager_env.commands.motion.motion_lib_cfg.motion_file=<path/to/robot_filtered> \
    ++manager_env.commands.motion.motion_lib_cfg.smpl_motion_file=<path/to/smpl_filtered>
```

For example, using the sample data from Hugging Face:

```bash
python gear_sonic/train_agent_trl.py \
    +exp=manager/universal_token/all_modes/sonic_release \
    num_envs=16 headless=True \
    ++manager_env.commands.motion.motion_lib_cfg.motion_file=sample_data/robot_filtered \
    ++manager_env.commands.motion.motion_lib_cfg.smpl_motion_file=sample_data/smpl_filtered
```

Or using the full dataset:

```bash
python gear_sonic/train_agent_trl.py \
    +exp=manager/universal_token/all_modes/sonic_release \
    num_envs=4096 headless=True \
    ++manager_env.commands.motion.motion_lib_cfg.motion_file=data/motion_lib_bones_seed/robot_filtered \
    ++manager_env.commands.motion.motion_lib_cfg.smpl_motion_file=data/smpl_filtered
```

### Finetuning from the released checkpoint

```bash
python gear_sonic/train_agent_trl.py \
    +exp=manager/universal_token/all_modes/sonic_release \
    +checkpoint=sonic_release/last.pt \
    num_envs=4096 headless=True \
    ++manager_env.commands.motion.motion_lib_cfg.motion_file=<path/to/robot_filtered> \
    ++manager_env.commands.motion.motion_lib_cfg.smpl_motion_file=<path/to/smpl_filtered>
```

### Multi-GPU and multi-node training

We recommend training with **64+ GPUs** for reasonable convergence times.
Single-node (8 GPU) training works but is significantly slower.

```bash
# Single node (8 GPUs)
accelerate launch --num_processes=8 gear_sonic/train_agent_trl.py \
    +exp=manager/universal_token/all_modes/sonic_release \
    num_envs=4096 headless=True

# Multi-node β€” use accelerate config for distributed setup
accelerate launch \
    --multi_gpu \
    --num_machines=8 \
    --num_processes=64 \
    --machine_rank=$MACHINE_RANK \
    --main_process_ip=$MASTER_ADDR \
    --main_process_port=$MASTER_PORT \
    gear_sonic/train_agent_trl.py \
    +exp=manager/universal_token/all_modes/sonic_release \
    num_envs=4096 headless=True
```

For multi-node setup, see the
[Accelerate distributed training guide](https://huggingface.co/docs/accelerate/usage_guides/deepspeed)
and
[multi-node launcher docs](https://huggingface.co/docs/accelerate/package_reference/cli#accelerate-launch).

### W&B logging

Enabled by default. Key overrides:

```bash
WANDB_MODE=offline python gear_sonic/train_agent_trl.py ...   # offline mode
    wandb.wandb_project=my_project wandb.wandb_entity=my_team  # custom project
    use_wandb=false                                             # disable entirely
```

### Local debug run

```bash
python gear_sonic/train_agent_trl.py \
    +exp=manager/universal_token/all_modes/sonic_release \
    num_envs=16 headless=False \
    ++algo.config.num_learning_iterations=100
```

## Monitoring

### Key metrics

| Metric | Good range | Description |
|--------|-----------|-------------|
| `rewards/total` | 3.0+ | Total reward |
| `rewards/anchor_pos_err` | < 0.15 | Root position tracking error (m) |
| `rewards/body_pos_err` | < 0.10 | Body position tracking error (m) |
| `throughput/fps` | ~4000+ | Training throughput |

### Checkpoints

Saved every 2000 steps to:

```
logs_rl/TRL_G1_Track/<experiment_name>-<timestamp>/
β”œβ”€β”€ model_step_002000.pt
β”œβ”€β”€ config.yaml
└── ...
```

## Evaluation

### Visualize reference motions

Replay motions to verify data quality before training:

```bash
python gear_sonic/train_agent_trl.py \
    +exp=manager/universal_token/all_modes/sonic_release \
    ++replay=True num_envs=4 headless=False
```

### Evaluate a checkpoint

Two eval modes: **metrics** (success rate, MPJPE) and **render** (video output).

For the released checkpoint, you must override motion paths since its
`config.yaml` has internal training paths. For your own checkpoints trained
with `sonic_release`, omit the motion overrides.

```bash
# --- Metrics ---
python gear_sonic/eval_agent_trl.py \
    +checkpoint=<path_to_checkpoint.pt> \
    +headless=True \
    ++eval_callbacks=im_eval \
    ++run_eval_loop=False \
    ++num_envs=128 \
    "+manager_env/terminations=tracking/eval" \
    "++manager_env.commands.motion.motion_lib_cfg.max_unique_motions=512"
```

```bash
# --- Render videos ---
python gear_sonic/eval_agent_trl.py \
    +checkpoint=<path_to_checkpoint.pt> \
    +headless=True \
    ++eval_callbacks=im_eval \
    ++run_eval_loop=False \
    ++num_envs=8 \
    ++manager_env.config.render_results=True \
    "++manager_env.config.save_rendering_dir=/tmp/renders" \
    ++manager_env.config.env_spacing=10.0 \
    "~manager_env/recorders=empty" "+manager_env/recorders=render"
```

For the **released checkpoint only**, append this override to either command
(its embedded config has internal training paths):

```bash
    "++manager_env.commands.motion.motion_lib_cfg.motion_file=data/motion_lib_bones_seed/robot_filtered"
```

Videos are saved as `000000.mp4`, `000001.mp4`, etc. in `save_rendering_dir`.

### Expected eval metrics

*Training rewards* (W&B `Episode_Reward/`):

| Metric | Converged | Description |
|--------|-----------|-------------|
| `tracking_vr_5point_local` | > 0.80 | 5-point tracking quality |
| `tracking_relative_body_pos` | > 0.44 | Upper-body position tracking |
| `tracking_anchor_pos` | > 0.14 | Root position tracking |
| `time_out` | > 0.90 | Episode completion rate |

*Eval metrics* (from `eval_agent_trl.py`):

| Metric | Converged | Description |
|--------|-----------|-------------|
| `success_rate` | > 0.97 | Motions tracked without early termination |
| `mpjpe_l` | < 30 mm | Local per-joint position error |
| `mpjpe_g` | < 200 mm | Global per-joint position error |

A well-converged policy reaches >0.98 success rate and <29 mm mpjpe_l after
100K iterations.

## ONNX Export

Export a trained checkpoint to ONNX for C++ deployment:

```bash
python gear_sonic/eval_agent_trl.py \
    +checkpoint=<path_to_checkpoint.pt> \
    +headless=True ++num_envs=1 \
    +export_onnx_only=true
```

For the released checkpoint, append the motion path overrides shown in the
eval section above.

Output (in `exported/` next to the checkpoint):

| File | Description |
|------|-------------|
| `*_smpl.onnx` | SMPL encoder + decoder (pose estimation input) |
| `*_g1.onnx` | G1 encoder + decoder (robot joint input) |
| `*_teleop.onnx` | Teleop encoder + decoder (VR tracking input) |
| `*_encoder.onnx` | All encoders combined |
| `*_decoder.onnx` | Decoder only |

Use the encoder+decoder pair matching your input modality. See
[deployment code reference](../references/deployment_code.md) for C++ details.

## Training with SOMA encoder

The `sonic_bones_seed` config adds a fourth SOMA encoder for BVH-derived
skeleton joint positions.

### SOMA data preparation

```bash
# Extract SOMA joints from BVH
python gear_sonic/data_process/extract_soma_joints_from_bvh.py \
    --input /path/to/bones_seed/bvh/ \
    --output data/motion_lib_bones_seed/soma \
    --fps 30 --num_workers 16 --skip_existing

# Filter to match robot data
python gear_sonic/data_process/filter_and_copy_bones_data.py \
    --source data/motion_lib_bones_seed/soma \
    --dest data/motion_lib_bones_seed/soma_filtered \
    --workers 16
```

### Training

Use multi-node training (64+ GPUs recommended):

```bash
accelerate launch \
    --multi_gpu --num_machines=8 --num_processes=64 \
    --machine_rank=$MACHINE_RANK \
    --main_process_ip=$MASTER_ADDR \
    --main_process_port=$MASTER_PORT \
    gear_sonic/train_agent_trl.py \
    +exp=manager/universal_token/all_modes/sonic_bones_seed \
    num_envs=4096 headless=True \
    ++manager_env.commands.motion.motion_lib_cfg.motion_file=data/motion_lib_bones_seed/robot_filtered \
    ++manager_env.commands.motion.motion_lib_cfg.smpl_motion_file=data/smpl_filtered \
    ++manager_env.commands.motion.motion_lib_cfg.soma_motion_file=data/motion_lib_bones_seed/soma_filtered
```

Data layout for 4-encoder training:

```
data/
β”œβ”€β”€ motion_lib_bones_seed/
β”‚   β”œβ”€β”€ robot_filtered/     # ~130K PKLs (G1 retargeted)
β”‚   └── soma_filtered/      # ~130K PKLs (SOMA skeleton)
└── smpl_filtered/          # ~131K PKLs (SMPL human)
```